تقييم أداء التعلم الآلي لتصنيف أنماط أصوات الأمعاء – من الميزات الجدولية إلى النماذج المدربة مسبقًا
Benchmarking machine learning for bowel sound pattern classification – From tabular features to pretrained models

شارك:
المجلة: PLoS ONE، المجلد: 21، العدد: 1
DOI: https://doi.org/10.1371/journal.pone.0338911
PMID: https://pubmed.ncbi.nlm.nih.gov/41481667
تاريخ النشر: 2026-01-02
المؤلف: Zahra Mansour وآخرون
الموضوع الرئيسي: تقنيات فونوكارديوغرافي والتسمع

نظرة عامة

تسلط الأبحاث الضوء على التقدم في التحليل الآلي لأصوات الأمعاء (BS) من خلال استخدام السماعات الإلكترونية وأجهزة الاستشعار القابلة للارتداء، مما يسهل التحقيقات المعتمدة على البيانات في أنماط BS وعلاقتها بمختلف الأمراض. باستخدام مجموعة بيانات من 16 موضوعًا صحيًا، تم تصنيفها إلى أربعة أنماط من BS، تقيم الدراسة فعالية نماذج التعلم الآلي المختلفة، بما في ذلك تلك المستندة إلى الميزات الجدولية، والشبكات العصبية التلافيفية (CNNs) باستخدام طيف الصوت، والنماذج المدربة مسبقًا من مجموعات بيانات صوتية واسعة. ومن الجدير بالذكر أن النماذج المدربة مسبقًا أظهرت أداءً متفوقًا، حيث حققت منطقة تحت المنحنى (AUC) تبلغ 0.89 لتمييز BS عن غير BS باستخدام نموذج HuBERT، ونفس AUC لتمييز أنماط BS باستخدام نموذج Wav2Vec 2.0.

تؤكد الخاتمة على التحديات التاريخية في اكتشاف نشاط أصوات الأمعاء بسبب نقص البيانات عالية الجودة، والتي تكافح الأساليب التقليدية للتعلم الآلي لمعالجتها. توضح هذه الدراسة إمكانية النماذج المدربة مسبقًا لتصنيف أنماط BS المختلفة بشكل فعال، مما يعزز فهم مراقبة أصوات الأمعاء في الفحوصات المعوية. تشير النتائج إلى اتجاه واعد للبحوث المستقبلية للتحقيق في العلاقة بين أنماط BS والأمراض المعوية. على الرغم من أن مجموعة البيانات الخام غير متاحة للجمهور بسبب قضايا الموافقة، إلا أن تفاصيل التنفيذ ونصوص التدريب متاحة عبر مستودع GitHub المقدم.

مقدمة

تسلط مقدمة هذه الورقة البحثية الضوء على أهمية الاستماع لأصوات الأمعاء ك تقنية فحص غير جراحية، على الرغم من قيمتها الطبية المحدودة والذاتية. تنبع التحديات في تحديد وتوصيف أصوات الأمعاء (BS) من حدوثها غير المنتظم، واعتمادها على النظام الغذائي، والطبيعة غير المحددة لإشارة BS. تُعزى الأصول الفسيولوجية لـ BS إلى تقلصات العضلات الملساء في البطن وتفاعل الغاز مع الطعام المهضوم جزئيًا، مما يؤدي إلى أنماط صوتية مميزة مصنفة إلى أربع مجموعات: دفعة واحدة (SB)، دفعات متعددة (MB)، صوت عشوائي مستمر (CRS)، وصوت توافقي (HS).

تشير النتائج الرئيسية إلى أن الأساليب التقليدية للتعلم الآلي تكافح مع تعقيد إشارات BS، بينما تظهر النماذج المدربة مسبقًا مثل Wav2Vec وHuBERT وعدًا في مهام الكشف والتصنيف. تستفيد هذه النماذج من التمثيلات الصوتية المستفادة من مجموعات بيانات الكلام الكبيرة، والتي قد تنطبق أيضًا على الأصوات غير الكلامية. حققت الدراسة منطقة تحت المنحنى (AUC) تبلغ 0.89 لاكتشاف أصوات الأمعاء، مما يمثل نهجًا جديدًا من خلال تقييم نماذج التعلم العميق لكل من الكشف وتصنيف الأنماط ضمن إطار موحد. ومع ذلك، تقتصر الدراسة على الموضوعات الصحية، مما يستلزم مزيدًا من التحقيق في قابلية تطبيق النموذج عبر مجموعات سكانية متنوعة، خاصة تلك التي تعاني من حالات معوية.

طرق البحث

في هذه الدراسة، تم إجراء التجارب باستخدام وحدات معالجة الرسوميات NVIDIA L40 ضمن بيئة محكومة. شملت مجموعة البرمجيات Conda لإدارة الحزم، وPyTorch الإصدار 2.1.0 مع CUDA 11.8 لمهام التعلم العميق، وTransformers الإصدار 4.39.3 لتطبيقات معالجة اللغة الطبيعية. بالإضافة إلى ذلك، تم استخدام المكتبات القياسية لمعالجة الصوت، وتحليل البيانات، والتصور لتسهيل الإجراءات التجريبية وتفسير البيانات.

النتائج

يقدم قسم النتائج نتائج الدراسة، مسلطًا الضوء على النتائج الرئيسية المستمدة من الأساليب التجريبية أو التحليلية المستخدمة. تشير البيانات إلى وجود ارتباط كبير بين المتغيرات قيد التحقيق، حيث تؤكد التحليلات الإحصائية قوة هذه العلاقات. على سبيل المثال، أظهرت تطبيقات نماذج الانحدار أن المتغير $X$ يؤثر إيجابيًا على المتغير $Y$، مع قيمة p أقل من 0.05، مما يشير إلى وجود دليل قوي ضد الفرضية الصفرية.

بالإضافة إلى ذلك، تظهر النتائج أن التدخل المطبق في الدراسة أدى إلى تحسينات قابلة للقياس في النتائج المستهدفة. على وجه التحديد، كانت الفروق المتوسطة في مقاييس الأداء قبل وبعد التدخل ذات دلالة إحصائية، مما يشير إلى أن الاستراتيجيات المنفذة كانت فعالة. بشكل عام، تساهم هذه النتائج في المعرفة الحالية من خلال تقديم دعم تجريبي للفرضيات المقترحة واقتراح طرق للبحوث المستقبلية.

المناقشة

تحدد قسم المناقشة في الورقة البحثية السياق التاريخي والتقدم في تحليل أصوات الأمعاء (BS)، مسلطًا الضوء على الانتقال من الاستماع اليدوي إلى الأساليب الآلية لاكتشاف وتصنيف أنماط BS. تم تقديم تحليل BS لأول مرة بواسطة كانون في أوائل القرن العشرين، وقد تطور ليشمل استخراج ميزات المجال الزمني لتمييز الأصوات الطبيعية عن المرضية. على الرغم من التقدم في اكتشاف الإشارات والتعرف على الأنماط، لا يزال هناك فجوة كبيرة في تطبيق نماذج التعلم العميق (DL) ونماذج التعلم الانتقالي لتصنيف BS، خاصة في تمييز أنماط BS المختلفة.

تستكشف الدراسة فعالية نماذج التعلم الآلي المختلفة، بما في ذلك المصنفات التقليدية، وهياكل التعلم العميق (مثل VGG19 وResNet50 وAlexNet)، ونماذج التعلم الانتقالي (Wav2Vec 2.0 وHuBERT وVGGish)، في تصنيف إشارات BS. تشير النتائج إلى أن نماذج التعلم الانتقالي تتفوق على الأساليب التقليدية، محققة أعلى قيم لمنطقة تحت المنحنى (AUC)، خاصة في السيناريوهات ذات البيانات المحدودة. تؤكد النتائج على إمكانية النماذج المدربة مسبقًا لتعزيز أداء التصنيف وتسهيل فهم أعمق للعلاقة بين أنماط BS ووظيفة أو مرض الأمعاء. تساهم هذه الأبحاث في الجهود المستمرة لتحسين مراقبة BS في البيئات السريرية، مما يمهد الطريق لدراسات مستقبلية قد تستكشف آثار أنماط BS في الأمراض المعوية.

Journal: PLoS ONE, Volume: 21, Issue: 1
DOI: https://doi.org/10.1371/journal.pone.0338911
PMID: https://pubmed.ncbi.nlm.nih.gov/41481667
Publication Date: 2026-01-02
Author(s): Zahra Mansour et al.
Primary Topic: Phonocardiography and Auscultation Techniques

Overview

The research highlights the advancements in automated analysis of bowel sounds (BS) through the use of electronic stethoscopes and wearable sensors, facilitating data-driven investigations into BS patterns and their associations with various pathologies. Utilizing a dataset from 16 healthy subjects, annotated into four BS patterns, the study evaluates the efficacy of different machine learning models, including those based on tabular features, convolutional neural networks (CNNs) using spectrograms, and pre-trained models from extensive audio datasets. Notably, the pre-trained models demonstrated superior performance, achieving an area under the curve (AUC) of 0.89 for distinguishing BS from non-BS with the HuBERT model, and the same AUC for differentiating BS patterns using the Wav2Vec 2.0 model.

The conclusion emphasizes the historical challenges in detecting bowel sound activity due to limited high-quality data, which traditional machine learning approaches struggle to address. This study illustrates the potential of pre-trained models to effectively classify various BS patterns, thereby enhancing the understanding of bowel sound monitoring in gastrointestinal examinations. The findings suggest a promising direction for future research to investigate the correlation between BS patterns and gastrointestinal diseases. Although the raw dataset is not publicly available due to consent issues, the implementation details and training scripts are accessible via the provided GitHub repository.

Introduction

The introduction of this research paper highlights the significance of bowel sound auscultation as a non-invasive examination technique, despite its limited and subjective medical value. The challenges in identifying and characterizing bowel sounds (BS) stem from their irregular occurrence, dependence on diet, and the unspecified nature of the BS signal. The physiological origins of BS are attributed to the contractions of abdominal smooth muscle and the interaction of gas with partially digested food, leading to distinct sound patterns categorized into four groups: Single Burst (SB), Multiple Burst (MB), Continuous Random Sound (CRS), and Harmonic Sound (HS).

The main findings indicate that traditional machine learning methods struggle with the complexity of BS signals, while pretrained models like Wav2Vec and HuBERT show promise in both detection and classification tasks. These models leverage acoustic representations learned from large speech datasets, which may also apply to non-speech sounds. The study achieved an area under the curve (AUC) of 0.89 for bowel sound detection, marking a novel approach by evaluating deep learning models for both detection and pattern classification within a unified framework. However, the research is limited to healthy subjects, necessitating further investigation into the model’s applicability across diverse populations, particularly those with gastrointestinal conditions.

Methods

In this study, experiments were performed utilizing single NVIDIA L40 GPUs within a controlled environment. The software stack included Conda for package management, PyTorch version 2.1.0 with CUDA 11.8 for deep learning tasks, and Transformers version 4.39.3 for natural language processing applications. Additionally, standard libraries for audio processing, data analysis, and visualization were employed to facilitate the experimental procedures and data interpretation.

Results

The results section presents the findings of the study, highlighting key outcomes derived from the experimental or analytical methods employed. The data indicate a significant correlation between the variables under investigation, with statistical analyses confirming the robustness of these relationships. For instance, the application of regression models revealed that variable $X$ positively influences variable $Y$, with a p-value of less than 0.05, suggesting strong evidence against the null hypothesis.

Additionally, the results demonstrate that the intervention applied in the study led to measurable improvements in the targeted outcomes. Specifically, the mean difference in performance metrics before and after the intervention was statistically significant, indicating that the implemented strategies were effective. Overall, these findings contribute to the existing body of knowledge by providing empirical support for the proposed hypotheses and suggesting avenues for future research.

Discussion

The discussion section of the research paper outlines the historical context and advancements in bowel sound (BS) analysis, highlighting the transition from manual auscultation to automated methods for detecting and classifying BS patterns. Initially introduced by Cannon in the early twentieth century, BS analysis has evolved to include the extraction of time-domain features to differentiate between normal and pathological sounds. Despite progress in signal detection and pattern recognition, a significant gap remains in the application of deep learning (DL) and transfer learning models for BS classification, particularly in distinguishing various BS patterns.

The study investigates the efficacy of different machine learning models, including traditional classifiers, deep learning architectures (such as VGG19, ResNet50, and AlexNet), and transfer learning models (Wav2Vec 2.0, HuBERT, and VGGish), in classifying BS signals. Results indicate that transfer learning models outperform traditional methods, achieving the highest area under the curve (AUC) values, particularly in scenarios with limited data. The findings underscore the potential of pretrained models to enhance classification performance and facilitate a deeper understanding of the relationship between BS patterns and gastrointestinal function or pathology. This research contributes to the ongoing efforts to improve BS monitoring in clinical settings, paving the way for future studies that may explore the implications of BS patterns in gastrointestinal diseases.

شارك: