DOI: https://doi.org/10.3390/electronics14020248
تاريخ النشر: 2025-01-09
المؤلف: Somaye Valizade Shayegh وآخرون
الموضوع الرئيسي: صحة وتطور الرضع
نظرة عامة
تتناول الدراسة القضية الملحة لمعدل وفيات حديثي الولادة، وخاصة في البيئات ذات الموارد المحدودة، من خلال استكشاف إمكانية استخدام أصوات بكاء المواليد كعلامات حيوية غير جراحية للتشخيص المبكر للحالات المهددة للحياة مثل الإنتان ومتلازمة ضيق التنفس (RDS). من خلال التركيز على أجزاء الزفير من إشارات البكاء، والتي تعتبر الأكثر معلوماتية، استخدم الباحثون نماذج التعلم الذاتي (SSL) – تحديدًا wav2vec 2.0 وWavLM وHuBERT – لاستخراج تمثيلات الميزات من الصوت الخام للبكاء. تلغي هذه الطريقة الحاجة إلى استخراج الميزات يدويًا وتلتقط بفعالية الأنماط المعقدة التي تشير إلى الإنتان وRDS.
نفذت الدراسة مصنفًا مع طبقة واحدة متصلة بالكامل فوق نماذج SSL لتصنيف حديثي الولادة إلى مجموعات صحية، إنتان، أو RDS. تم إجراء تحسين المعلمات باستخدام استراتيجيتين لمعدل التعلم: خطية وتبريد. أظهرت النتائج أن استراتيجية التبريد تفوقت باستمرار على النهج الخطي، حيث حقق wav2vec 2.0 أعلى دقة تصنيف تبلغ حوالي 90% (89.76%). تشير هذه النتائج إلى أن دمج هذه المنهجية في أنظمة تشخيص بكاء حديثي الولادة (NCDSs) يمكن أن يعزز بشكل كبير قدرة الطاقم الطبي على تحديد حديثي الولادة المرضى بشكل حرج، مما يسهل التدخلات في الوقت المناسب ويحسن معدلات بقاء حديثي الولادة.
مقدمة
في عام 2023، كان معدل وفيات الأطفال دون سن الخامسة عالميًا حوالي 38 وفاة لكل 1000 ولادة حية، مع تقدير وفيات حديثي الولادة بـ 17 لكل 1000. تحدث وفيات حديثي الولادة، التي بلغت حوالي 2.3 مليون في عام 2022، بشكل رئيسي في البلدان ذات الدخل المنخفض حيث الوصول إلى الرعاية الصحية محدود. يعد التشخيص المبكر للحالات الحرجة مثل الإنتان ومتلازمة ضيق التنفس (RDS) أمرًا ضروريًا، حيث يمكن أن تؤدي الحالات غير المعالجة إلى مضاعفات شديدة ومعدلات وفيات مرتفعة. يمثل الإنتان 20% إلى 36% من وفيات حديثي الولادة، بينما يمكن أن تتجاوز وفيات RDS 30% عندما تتأخر التدخلات. يتم التأكيد على الحاجة الملحة لأدوات تشخيصية متاحة في البيئات ذات الموارد المحدودة، خاصة بالنظر إلى الملاحظات التاريخية التي تشير إلى أن بكاء حديثي الولادة يمكن أن يدل على مشاكل صحية كامنة.
تهدف هذه الدراسة إلى تطوير نظام تحليل بكاء غير جراحي (NCDS) للكشف المبكر عن الإنتان وRDS من خلال تحليل الأنماط الصوتية في بكاء الرضع. باستخدام مجموعة بيانات خاصة تزيد من حجم العينة بمقدار 2.5 مرة وتضمن تمثيلًا متوازنًا عبر الحالات (RDS، صحي، وإنتان)، تستخدم الأبحاث نماذج التعلم الذاتي (SSL) لالتقاط الأنماط المعقدة للبكاء مباشرة من الصوت الخام. تلغي هذه الطريقة الحاجة إلى هندسة الميزات يدويًا، مما يعزز دقة التمييز بين حالات الصحة. تم هيكلة الورقة لتشمل مراجعة الأدبيات، منهجية مفصلة، نتائج تجريبية، ومناقشة للنتائج، مما يساهم في تطوير أنظمة NCDSs المتقدمة التي يمكن أن تحسن نتائج صحة حديثي الولادة.
طرق
تستخدم المنهجية الموضحة في هذا البحث الموجات الخام كميزات إدخال لمصنف الشبكة العصبية ونظام القرار (NCDS)، مستفيدة من نماذج التعلم الذاتي لاستخراج الميزات وطبقة متصلة بالكامل للتصنيف. استخدم الإعداد التجريبي خادم Narval، المجهز بوحدات معالجة الرسوميات Nvidia A100 ووحدات المعالجة المركزية AMD Milan، مع تخصيص كل وظيفة لوحدة معالجة رسومات واحدة تحتوي على 25 جيجابايت من الذاكرة. تم إجراء معالجة البيانات المسبقة باستخدام مكتبة TorchAudio وحزمة Librosa، بينما تم تسهيل ضبط النماذج بواسطة مكتبة Transformers وSpeechBrain وPyTorch.
تركز النتائج التجريبية على أداء النماذج المضبوطة في التمييز بين ثلاث فئات مرضية: RDS، الإنتان، وصحي. اختبرت الدراسة ثلاثة نماذج تستخدم استراتيجيتين لمعدل التعلم – خطية وتبريد – لتحسين دقة التصنيف. تشير النتائج إلى فعالية المنهجية المقترحة في التمييز بدقة بين هذه الفئات، مما يبرز إمكانيات التعلم الذاتي في التشخيص الطبي.
نتائج
يقدم قسم النتائج نتائج من التجارب مع النماذج المضبوطة – wav2vec 2.0 وHuBERT وWavLM+ – باستخدام استراتيجيات معدل التعلم لكل من التبريد والخطية. بالنسبة لمعدل التعلم بالتبريد، تضمنت التكوينات المثلى 11 إلى 12 دورة، وحجم دفعة ثابت قدره 8، وانخفاض وزن قدره $1 \times 10^{-4}$. تم تحديد معدلات التعلم الذاتي (SSL) كـ $1.7 \times 10^{-5}$ لـ wav2vec 2.0، و$3.6 \times 10^{-5}$ لـ HuBERT، و$1.9 \times 10^{-5}$ لـ WavLM+، مع كون معدلات الشبكة متعددة الطبقات (MLP) أعلى بعشر مرات. أظهرت أداء التدريب والتحقق، الموضح في الشكل 4، تحسينات كبيرة في معدلات الخسارة والخطأ عبر النماذج، حيث حقق wav2vec 2.0 تقليل خطأ التحقق من 0.473 إلى 0.224، وWavLM+ من 0.436 إلى 0.198، وHuBERT من 0.365 إلى 0.221.
في المقابل، أسفرت استراتيجية معدل التعلم الخطي عن دورات مثلى قدرها 9 و12، مع الحفاظ على حجم دفعة قدره 8. تم ضبط معدلات التعلم الأولية إلى $2 \times 10^{-5}$ لـ wav2vec 2.0، و$2.1 \times 10^{-5}$ لـ HuBERT، و$2.7 \times 10^{-5}$ لـ WavLM+، متقاربة إلى معدل نهائي قدره $1 \times 10^{-7}$. يوضح الشكل 6 خسارة التدريب ومعدلات خطأ التحقق، مع إظهار wav2vec 2.0 تقليل الخسارة من 1.13 إلى 0.0061 وخطأ التحقق من 0.667 إلى 0.229. كما أظهرت WavLM+ وHuBERT تحسينات كبيرة، محققة أداءً مثاليًا في دورات مختلفة. يتم تلخيص مقاييس الأداء، بما في ذلك الدقة، الدقة، الاسترجاع، ودرجات F1، في الجداول 5 و7، مما يشير إلى قدرات التعميم القوية للنماذج في تصنيف بكاء الأطفال إلى حالات صحية، إنتان، وRDS، مع دقة تتراوح من 88.33% إلى 89.76%.
مناقشة
في قسم المناقشة من ورقة البحث، يبرز المؤلفون التقدم والتحديات في تطبيقات التعلم الآلي لتصنيف بكاء الرضع، وخاصة في التمييز بين الحالات الصحية والمرضية. حققت خوارزميات التعلم الآلي معدلات دقة ملحوظة، حيث وصلت بعض النماذج إلى 80.56% في التعرف على البكاء المرتبط بالاحتياجات الأساسية، متفوقة بشكل كبير على المستمعين البشر. ومع ذلك، لا يزال تصنيف الأمراض معقدًا بسبب الأنماط الصوتية المعقدة المرتبطة بالبكاء المرضي. يؤكد المؤلفون أن مهام التصنيف الثنائي، مثل التمييز بين البكاء الصحي وغير الصحي، قد أسفرت عن نتائج أفضل من تصنيف الأمراض متعددة الفئات، والذي غالبًا ما يواجه صعوبات مع مجموعات البيانات المحدودة وتقنيات استخراج الميزات.
تحدد الورقة ثلاثة نهج رئيسية في تصنيف الأمراض بناءً على البكاء: التمييز بين البكاء الصحي وغير الصحي، تحديد الأمراض المحددة، والتوسع إلى تصنيف متعدد الفئات. استخدمت دراسات مختلفة طرق استخراج ميزات مختلفة، بما في ذلك معاملات cepstral بتردد ميل (MFCCs) وتقنيات التعلم العميق، لتعزيز دقة التصنيف. على الرغم من هذه الجهود، لا تزال هناك فجوات كبيرة في الأدبيات، خاصة فيما يتعلق بتصنيف الأمراض متعددة الفئات ذات معدلات الوفيات العالية والحاجة إلى مجموعات بيانات أكثر شمولاً. يساهم المؤلفون في هذا المجال من خلال توسيع مجموعة بياناتهم بشكل كبير وتقديم مجموعات ميزات صوتية عميقة قوية مستمدة من الصوت الخام، مستفيدين من تقنيات التعلم الذاتي. تهدف هذه الطريقة إلى تحسين قدرة النموذج على التعميم عبر سيناريوهات سريرية متنوعة، مما يعزز في النهاية تطوير نظام تصنيف بكاء حديثي الولادة متعدد الفئات.
DOI: https://doi.org/10.3390/electronics14020248
Publication Date: 2025-01-09
Author(s): Somaye Valizade Shayegh et al.
Primary Topic: Infant Health and Development
Overview
The study addresses the pressing issue of neonatal mortality, particularly in resource-limited settings, by exploring the potential of newborn cry sounds as non-invasive biomarkers for early diagnosis of life-threatening conditions such as Sepsis and Respiratory Distress Syndrome (RDS). By focusing on the expiratory segments of cry signals, which are deemed most informative, the researchers employed Self-Supervised Learning (SSL) models—specifically wav2vec 2.0, WavLM, and HuBERT—to extract feature representations from raw cry audio. This approach negates the need for manual feature extraction and effectively captures the complex patterns indicative of sepsis and RDS.
The study implemented a classifier with a single fully connected layer atop the SSL models to categorize newborns into Healthy, Sepsis, or RDS groups. Hyperparameter optimization was conducted using two learning rate strategies: linear and annealing. The results indicated that the annealing strategy consistently outperformed the linear approach, with wav2vec 2.0 achieving the highest classification accuracy of approximately 90% (89.76%). These findings suggest that integrating this methodology into Newborn Cry Diagnosis Systems (NCDSs) could significantly enhance the ability of medical personnel to identify critically ill newborns, thereby facilitating timely interventions and improving neonatal survival rates.
Introduction
In 2023, the global under-five child mortality rate was approximately 38 deaths per 1000 live births, with neonatal mortality estimated at 17 per 1000. Neonatal deaths, totaling around 2.3 million in 2022, predominantly occur in low-income countries where access to healthcare is limited. Early diagnosis of critical conditions such as sepsis and respiratory distress syndrome (RDS) is essential, as untreated cases can lead to severe complications and high mortality rates. Sepsis accounts for 20% to 36% of neonatal deaths, while RDS mortality can exceed 30% when interventions are delayed. The urgent need for accessible diagnostic tools in low-resource settings is emphasized, particularly given historical observations that neonatal cries can indicate underlying health issues.
This study aims to develop a non-invasive cry analysis system (NCDS) for the early detection of sepsis and RDS by analyzing acoustic patterns in infant cries. Utilizing a private dataset that expands the sample size by 2.5 times and ensures balanced representation across conditions (RDS, Healthy, and Sepsis), the research employs self-supervised learning (SSL) models to capture complex cry patterns directly from raw audio. This method eliminates the need for manual feature engineering, enhancing the accuracy of distinguishing between health conditions. The paper is structured to include a literature review, detailed methodology, experimental results, and a discussion of findings, ultimately contributing to the development of advanced NCDSs that can improve neonatal health outcomes.
Methods
The methodology outlined in this research employs raw waveforms as input features for the Neural Classifier and Decision System (NCDS), leveraging self-supervised learning models for feature extraction and a fully connected layer for classification. The experimental setup utilized the Narval server, equipped with Nvidia A100 GPUs and AMD Milan CPUs, with each job allocated one GPU featuring 25 GB of memory. Data pre-processing was conducted using the TorchAudio library and the Librosa package, while the fine-tuning of models was facilitated by the Transformers library, SpeechBrain, and PyTorch.
The experimental results focus on the performance of fine-tuned models in differentiating between three pathological classes: RDS, Sepsis, and Healthy. The study tested three models employing two learning rate strategies—linear and annealing—to optimize classification accuracy. The findings indicate the effectiveness of the proposed methodology in accurately distinguishing between these classes, highlighting the potential of self-supervised learning in medical diagnostics.
Results
The results section presents findings from experiments with fine-tuned models—wav2vec 2.0, HuBERT, and WavLM+—utilizing both annealing and linear learning rate strategies. For the annealing learning rate, optimal configurations included 11 to 12 epochs, a consistent batch size of 8, and a weight decay of $1 \times 10^{-4}$. The self-supervised learning (SSL) rates were determined as $1.7 \times 10^{-5}$ for wav2vec 2.0, $3.6 \times 10^{-5}$ for HuBERT, and $1.9 \times 10^{-5}$ for WavLM+, with multi-layer perceptron (MLP) rates being ten times higher. Training and validation performance, illustrated in Figure 4, showed significant improvements in loss and error rates across models, with wav2vec 2.0 achieving a validation error reduction from 0.473 to 0.224, WavLM+ from 0.436 to 0.198, and HuBERT from 0.365 to 0.221.
In contrast, the linear learning rate strategy yielded optimal epochs of 9 and 12, maintaining a batch size of 8. Initial learning rates were fine-tuned to $2 \times 10^{-5}$ for wav2vec 2.0, $2.1 \times 10^{-5}$ for HuBERT, and $2.7 \times 10^{-5}$ for WavLM+, converging to a final rate of $1 \times 10^{-7}$. Figure 6 depicts the training loss and validation error rates, with wav2vec 2.0 showing a loss reduction from 1.13 to 0.0061 and validation error from 0.667 to 0.229. WavLM+ and HuBERT also demonstrated substantial improvements, achieving optimal performance at various epochs. Performance metrics, including accuracy, precision, recall, and F1 scores, are summarized in Tables 5 and 7, indicating strong generalization capabilities of the models in classifying baby cries into Healthy, Sepsis, and RDS conditions, with accuracies ranging from 88.33% to 89.76%.
Discussion
In the discussion section of the research paper, the authors highlight the advancements and challenges in machine learning applications for classifying infant cries, particularly in distinguishing between healthy and pathological states. Machine learning algorithms have achieved notable accuracy rates, with some models reaching up to 80.56% in recognizing cries related to basic needs, significantly outperforming human listeners. However, the classification of diseases remains complex due to the intricate acoustic patterns associated with pathological cries. The authors emphasize that binary classification tasks, such as differentiating between healthy and unhealthy cries, have yielded better results than multiclass disease classification, which often struggles with limited datasets and feature extraction techniques.
The paper identifies three primary approaches in cry-based disease classification: distinguishing healthy from unhealthy cries, identifying specific diseases, and expanding to multiclass classification. Various studies have employed different feature extraction methods, including Mel-frequency cepstral coefficients (MFCCs) and deep learning techniques, to enhance classification accuracy. Despite these efforts, significant gaps remain in the literature, particularly regarding multiclass classification for diseases with high mortality rates and the need for more comprehensive datasets. The authors contribute to this field by expanding their dataset significantly and introducing robust deep audio feature sets derived from raw audio, leveraging self-supervised learning techniques. This approach aims to improve the model’s ability to generalize across diverse clinical scenarios, ultimately advancing the development of a multiclass neonatal cry classification system.
