DOI: https://doi.org/10.3389/fdgth.2026.1771281
PMID: https://pubmed.ncbi.nlm.nih.gov/41836340
تاريخ النشر: 2026-02-27
المؤلف: Abishek Ravichandran وآخرون
الموضوع الرئيسي: اضطرابات الصوت والكلام
نظرة عامة
تقدم هذه الورقة البحثية إطار عمل للتعلم العميق متعدد الأنماط للكشف المبكر عن مرض باركنسون (PD)، مع معالجة القيود التي تواجه الأساليب الأحادية التي تعتمد على التقييمات السريرية الذاتية. يدمج النموذج المقترح بيانات الكتابة اليدوية، والمشي، والكلام من خلال استراتيجية دمج الميزات المبكرة، مستخدمًا الشبكات العصبية العميقة لاستخراج الميزات ومصنف XGBoost للتنبؤات النهائية. يحقق الإطار دقة تصل إلى 92%، متجاوزًا الأنماط الفردية، ويظهر مقاييس أداء قوية، بما في ذلك درجة F1 الماكرو 0.89 ومنطقة تحت منحنى ROC (AUC) 0.95. يعزز استخدام تقنيات الذكاء الاصطناعي القابلة للتفسير، مثل SHapley Additive exPlanations (SHAP) وGradient-weighted Class Activation Mapping (Grad-CAM)، من قابلية تفسير النموذج من خلال تحديد الميزات الرئيسية التي تساهم في التنبؤ بـ PD.
تؤكد الخاتمة على الأداء المتفوق للنموذج مقارنة بالأسس الأحادية، خاصة في التعامل مع البيانات المزعجة أو غير المكتملة. يتم اقتراح تحسينات مستقبلية، بما في ذلك تطوير نماذج طولية لتتبع تقدم الأعراض، ودمج مجموعات بيانات متنوعة من العالم الحقيقي، وتطبيق تقنيات دمج متعددة الأنماط متقدمة. بالإضافة إلى ذلك، تدعو الورقة إلى دمج التعلم الفيدرالي الذي يحافظ على الخصوصية، وتقدير فترات الثقة، وإمكانية التنفيذ في الوقت الحقيقي على الأجهزة القابلة للارتداء. تهدف هذه التطورات إلى تحويل النموذج من نموذج بحثي إلى نظام دعم قرار قابل للنشر سريريًا، مما يعزز في النهاية قدرات الفحص المبكر والمراقبة لـ PD.
مقدمة
تتناول مقدمة هذه الورقة البحثية التحديات المرتبطة بتشخيص مرض باركنسون (PD)، وهو اضطراب عصبي تنكسي شائع يتميز بأعراض حركية وغير حركية. تعتمد طرق التشخيص التقليدية بشكل كبير على التقييمات السريرية الذاتية، مما قد يؤدي إلى تباين وقد لا تلتقط بفعالية مظاهر المرض المبكرة. لتعزيز دقة التشخيص، تسلط الورقة الضوء على إمكانيات أنظمة الذكاء الاصطناعي (AI) التي تحلل العلامات الحيوية الرقمية من الكلام، والمشي، والكتابة اليدوية. أظهرت التطورات الأخيرة في التعلم العميق نتائج واعدة في الأساليب الأحادية، محققة معدلات دقة عالية (تصل إلى 99% للكلام و98% للكتابة اليدوية). ومع ذلك، تواجه هذه الأنظمة الأحادية قيودًا كبيرة، بما في ذلك القابلية للتأثر بالعوامل البيئية، وجودة المستشعر، ونقص القابلية للتعميم.
لمعالجة هذه العيوب، تقترح الورقة إطار عمل للتعلم العميق متعدد الأنماط يدمج بيانات الكلام، والمشي، والكتابة اليدوية، مما يسمح بتحليل أكثر قوة وشمولية لـ PD. يستخدم هذا النظام دمج الميزات المبكرة لوزن موثوقية كل نمط بشكل ديناميكي، مما يعزز المرونة ضد الضوضاء والبيانات المفقودة. بالإضافة إلى ذلك، يهدف دمج تقنيات الذكاء الاصطناعي القابلة للتفسير، مثل SHapley Additive exPlanations (SHAP) وGradient-weighted Class Activation Mapping (Grad-CAM)، إلى تحسين قابلية تفسير تنبؤات النموذج، مما يعزز ثقة وفهم الأطباء. يتم اختبار الإطار المقترح بدقة على مجموعات بيانات مرجعية، مما يظهر دقة تشخيصية واستقرارًا متفوقين مقارنة بالنماذج الأحادية وغير التكيفية. في النهاية، تهدف هذه الأبحاث إلى سد الفجوة بين تطبيقات الذكاء الاصطناعي التجريبية والممارسة السريرية، مقدمةً حلاً موثوقًا، وقابلًا للتفسير، وقابلًا للتوسع للكشف عن PD.
الطرق
تدمج المنهجية المقدمة في هذا البحث ثلاثة أنماط—المشي، والصوت، والكتابة اليدوية—في أداة تشخيصية متعددة الأنماط لمرض باركنسون (PD). يستخدم كل نمط خط أنابيب متخصص لاستخراج الميزات: يتم تحليل إشارات المشي باستخدام الالتواءات الزمنية مع التمددات والترميز التلقائي، ويتم معالجة بيانات الصوت باستخدام EfficientNet-B0 مع الاستفادة من طيف log-Mel، ويتم فحص صور الكتابة اليدوية من خلال ResNet-50. تلتقط هذه الطريقة العلامات الحيوية الخاصة بكل نمط مثل شذوذ الخطوات، والأخطاء الصوتية، والشذوذ المرتبط بالاهتزاز. يتم تطبيع الميزات المستخرجة، وضغطها عبر التجميع المتوسط العالمي، وتحويلها إلى متجهات كامنة لتسهيل المقارنة عبر المجالات. يتم تعزيز قابلية تفسير النموذج من خلال تقنيات مثل SHAP، وIntegrated Gradients، وGrad-CAM، مما يسمح للأطباء بفهم عملية التنبؤ. يتم معالجة التمثيل متعدد الأنماط المدمج بعد ذلك بواسطة XGBoost، مما يحسن دقة التصنيف ويوفر رؤى ذات صلة سريريًا حول الأعراض المتنوعة لـ PD.
تستخدم الدراسة ثلاث مجموعات بيانات مرجعية متاحة للجمهور للكتابة اليدوية، والكلام، والمشي، مما يضمن عدم وجود تداخل على مستوى الموضوع عبر الأنماط لتعكس السيناريوهات السريرية الواقعية. تتكون مجموعة بيانات الكتابة اليدوية من 3,264 عينة رسم حلزوني من مرضى PD وأشخاص أصحاء، مما يلتقط العجز الحركي الدقيق الذي يدل على تدهور العضلات العصبية. يتم استخدام استراتيجية تقسيم البيانات على مستوى الموضوع للتجارب الأحادية، مع تخصيص 70% من الموضوعات للتدريب، و15% للتحقق، و15% للاختبار، باستخدام مخطط التحقق المتقاطع المنظم بخمسة أضعاف للحفاظ على تمثيل متوازن. بالنسبة للتجارب متعددة الأنماط، يتم محاذاة تضمينات الميزات على مستوى الميزات، وتستخدم التمثيلات الكامنة المدمجة للدمج، مع تطبيق تحقق متقاطع متسق عبر جميع النماذج لضمان مقارنة عادلة. يتم تقديم ملخص شامل لخصائص مجموعة البيانات في الجدول 2.
النتائج
تشير نتائج الدراسة إلى أن إطار الكشف عن مرض باركنسون (PD) متعدد الأنماط المقترح، الذي يدمج الكتابة اليدوية، والمشي، والكلام من خلال نهج دمج الميزات المبكرة، من المتوقع أن يتفوق على الأنظمة الأحادية، خاصة في السيناريوهات التي قد يتعرض فيها أحد الأنماط للخطر. يقوم الإطار بضبط الأوزان المعينة لكل نمط بشكل ديناميكي بناءً على موثوقيتها، مما يعزز دقة التشخيص. من المخطط إجراء تحليلات مقارنة لدعم الأداء المتفوق للنهج متعدد الأنماط مقابل الأسس الأحادية وغير التكيفية. بالإضافة إلى ذلك، ستساعد تقنيات الذكاء الاصطناعي القابلة للتفسير (XAI)، مثل SHAP وGrad-CAM، في توفير رؤى حول عمليات اتخاذ القرار للنموذج، مما يربط التنبؤات بميزات محددة مثل اهتزاز الكتابة اليدوية أو أنماط الكلام.
تكشف النتائج أن نموذج ResNet50 القائم على الكتابة اليدوية يظهر خصوصية وحساسية عالية في التمييز بين الأفراد الأصحاء ومرضى PD، بدقة تتراوح بين 88%-89% بعد 50 دورة. يلتقط النموذج بفعالية العلامات الحيوية للكتابة اليدوية، بما في ذلك الانحرافات الناتجة عن الاهتزاز، مما يؤكد قوة النمط في الكشف عن PD. يظهر نموذج الكلام أيضًا تعلمًا كبيرًا للعلامات الحيوية الصوتية، بينما يكشف تحليل المشي عن علاقات قوية بين قياسات زمنية وقوة متنوعة، مما يبرز الاعتماد المتبادل لمعايير المشي. تشير أداء الترميز التلقائي إلى تعلم ناجح لتمثيلات إشارات المشي، مع أخطاء إعادة بناء منخفضة، مما يشير إلى أن التضمينات مناسبة لمهام التصنيف اللاحقة. بشكل عام، يظهر الإطار متعدد الأنماط إمكانيات واعدة للكشف الدقيق والموثوق عن PD، مدعومًا بديناميات تحسين مستقرة ومعدلات دقة تحقق عالية.
المناقشة
تسلط قسم المناقشة في الورقة الضوء على التحديات في تشخيص مرض باركنسون (PD) بسبب تباين الأعراض والاعتماد على التقييمات السريرية الذاتية. تفتقر الطرق الحالية للذكاء الاصطناعي الأحادية، على الرغم من دقتها في مجالات معينة مثل تحليل الكلام أو المشي، إلى القابلية للتعميم وقابلية التفسير، مما يجعلها أقل ملاءمة للتطبيق السريري. يؤكد المؤلفون على الحاجة إلى إطار قوي متعدد الأنماط يدمج علامات PD الحيوية المختلفة، يمكنه التعامل مع البيانات المزعجة، ويوفر رؤى قابلة للتفسير لدعم اتخاذ القرار السريري.
تستعرض الورقة عدة نهج مبتكرة للكشف المبكر عن PD، مع التركيز على تحليل إشارات الكلام، وديناميات المشي، والكتابة اليدوية. بشكل ملحوظ، أظهرت تقنيات التعلم العميق، بما في ذلك الشبكات العصبية التلافيفية (CNNs) والنماذج الهجينة، معدلات دقة عالية (تصل إلى 99.67%) في التمييز بين مرضى PD في مراحل مبكرة والأفراد الأصحاء. يدعو المؤلفون إلى أنظمة متعددة الأنماط تجمع بين أنواع بيانات مختلفة، مثل الكلام، والمشي، والكتابة اليدوية، لتعزيز دقة التشخيص وموثوقيته. كما يتناولون قابلية تفسير نماذج الذكاء الاصطناعي من خلال تقنيات مثل Grad-CAM وSHAP، التي توضح عمليات اتخاذ القرار لهذه الأنظمة، مما يعزز الثقة بين الأطباء. بشكل عام، تؤكد النتائج على إمكانيات دمج مصادر بيانات متنوعة وتقنيات التعلم الآلي المتقدمة لتحسين الكشف المبكر عن PD والنتائج السريرية.
القيود
تعترف الدراسة بعدة قيود تتعلق بالإطار متعدد الأنماط المقترح. على الرغم من أنه يظهر قوة محسنة مقارنة بالطرق الأحادية عند تقييمه على مجموعات بيانات مرجعية استعادية، إلا أن قابليته للتكيف مقيدة بدمج الميزات على مستوى الميزات. علاوة على ذلك، لم يتم التحقق من صحة الإطار بعد في بيئات سريرية ديناميكية في العالم الحقيقي، مما قد يؤثر على قابليته للتطبيق العملي وفعاليته في سيناريوهات سريرية متنوعة.
DOI: https://doi.org/10.3389/fdgth.2026.1771281
PMID: https://pubmed.ncbi.nlm.nih.gov/41836340
Publication Date: 2026-02-27
Author(s): Abishek Ravichandran et al.
Primary Topic: Voice and Speech Disorders
Overview
This research paper presents a multimodal deep learning framework for the early detection of Parkinson’s disease (PD), addressing the limitations of unimodal approaches that rely on subjective clinical assessments. The proposed model integrates handwriting, gait, and speech data through an early feature fusion strategy, employing deep neural networks for feature extraction and an XGBoost classifier for final predictions. The framework achieves an accuracy of 92%, surpassing individual modalities, and demonstrates strong performance metrics, including a macro F1-score of 0.89 and an area under the ROC curve (AUC) of 0.95. The use of explainable AI techniques, such as SHapley Additive exPlanations (SHAP) and Gradient-weighted Class Activation Mapping (Grad-CAM), enhances the model’s interpretability by identifying key features contributing to PD prediction.
The conclusion emphasizes the model’s superior performance compared to unimodal baselines, particularly in handling noisy or incomplete data. Future enhancements are suggested, including the development of longitudinal models to track symptom progression, the incorporation of diverse real-world datasets, and the application of advanced multimodal fusion techniques. Additionally, the paper advocates for the integration of privacy-preserving federated learning, confidence interval estimation, and the potential for real-time implementation on wearable devices. These advancements aim to transition the model from a research prototype to a clinically deployable decision support system, ultimately enhancing early PD screening and monitoring capabilities.
Introduction
The introduction of this research paper addresses the challenges associated with diagnosing Parkinson’s disease (PD), a prevalent neurodegenerative disorder characterized by both motor and non-motor symptoms. Traditional diagnostic methods rely heavily on subjective clinical assessments, which can lead to variability and may not effectively capture early disease manifestations. To enhance diagnostic accuracy, the paper highlights the potential of artificial intelligence (AI) systems that analyze digital biomarkers from speech, gait, and handwriting. Recent advancements in deep learning have shown promising results in unimodal approaches, achieving high accuracy rates (up to 99% for speech and 98% for handwriting). However, these unimodal systems face significant limitations, including susceptibility to environmental factors, sensor quality, and lack of generalizability.
To address these shortcomings, the paper proposes a multimodal deep learning framework that integrates speech, gait, and handwriting data, allowing for a more robust and comprehensive analysis of PD. This system employs early feature fusion to dynamically weight the reliability of each modality, enhancing resilience against noise and missing data. Additionally, the incorporation of explainable AI techniques, such as SHapley Additive exPlanations (SHAP) and Gradient-weighted Class Activation Mapping (Grad-CAM), aims to improve the interpretability of the model’s predictions, fostering clinician trust and understanding. The proposed framework is rigorously tested on benchmark datasets, demonstrating superior diagnostic accuracy and stability compared to unimodal and non-adaptive models. Ultimately, this research aims to bridge the gap between experimental AI applications and clinical practice, offering a reliable, interpretable, and scalable solution for PD detection.
Methods
The methodology presented in this research integrates three modalities—gait, voice, and handwriting—into a multimodal diagnostic tool for Parkinson’s Disease (PD). Each modality employs a specialized feature extraction pipeline: gait signals are analyzed using temporal convolutions with dilations and autoencoding, voice data are processed with EfficientNet-B0 utilizing log-Mel spectrograms, and handwriting images are examined through ResNet-50. This approach captures modality-specific biomarkers such as stride anomalies, phonatory errors, and tremor-related anomalies. The extracted features are normalized, compressed via global average pooling, and transformed into latent vectors for cross-domain comparability. The model’s interpretability is enhanced through techniques like SHAP, Integrated Gradients, and Grad-CAM, allowing clinicians to understand the prediction process. The fused multimodal representation is then processed by XGBoost, optimizing classification accuracy and providing clinically relevant insights into the heterogeneous symptoms of PD.
The study employs three publicly available benchmark datasets for handwriting, speech, and gait, ensuring no subject-level overlap across modalities to reflect realistic clinical scenarios. The handwriting dataset consists of 3,264 spiral drawing samples from PD patients and healthy controls, capturing fine motor impairments indicative of neuromuscular degradation. A subject-wise data splitting strategy is utilized for unimodal experiments, with 70% of subjects for training, 15% for validation, and 15% for testing, using a fivefold stratified cross-validation scheme to maintain balanced representation. For multimodal experiments, feature embeddings are aligned at the feature level, and concatenated latent representations are used for fusion, with consistent cross-validation applied across all models to ensure fair comparison. A comprehensive summary of dataset characteristics is provided in Table 2.
Results
The results of the study indicate that the proposed multimodal Parkinson’s disease (PD) detection framework, which integrates handwriting, gait, and speech modalities through an early feature fusion approach, is expected to outperform unimodal systems, particularly in scenarios where one modality may be compromised. The framework dynamically adjusts the weights assigned to each modality based on their reliability, thereby enhancing diagnostic accuracy. Comparative analyses are planned to substantiate the superior performance of the multimodal approach against unimodal and non-adaptive baselines. Additionally, explainable artificial intelligence (XAI) techniques, such as SHAP and Grad-CAM, will facilitate insights into the model’s decision-making processes, linking predictions to specific features like handwriting tremors or speech patterns.
The findings reveal that the handwriting-based ResNet50 model demonstrates high specificity and sensitivity in distinguishing between healthy individuals and PD patients, with an accuracy of 88%-89% after 50 epochs. The model effectively captures handwriting biomarkers, including tremor-induced deviations, confirming the modality’s robustness in PD detection. The speech model also shows significant learning of vocal biomarkers, while gait analysis reveals strong correlations among various temporal and force-based measurements, underscoring the interdependence of gait parameters. The autoencoder’s performance indicates successful learning of gait signal representations, with low reconstruction errors, suggesting the embeddings are suitable for subsequent classification tasks. Overall, the multimodal framework exhibits promising potential for accurate and reliable PD detection, supported by stable optimization dynamics and high validation accuracy rates.
Discussion
The discussion section of the paper highlights the challenges in diagnosing Parkinson’s Disease (PD) due to symptom heterogeneity and reliance on subjective clinical ratings. Current unimodal AI methods, while accurate in specific domains such as speech or gait analysis, lack generalizability and interpretability, making them less suitable for clinical application. The authors emphasize the need for a robust, multimodal framework that integrates various PD biomarkers, can handle noisy data, and provides interpretable insights to support clinical decision-making.
The paper reviews several innovative approaches to early PD detection, focusing on the analysis of speech signals, gait dynamics, and handwriting. Notably, deep learning techniques, including convolutional neural networks (CNNs) and hybrid models, have demonstrated high accuracy rates (up to 99.67%) in distinguishing between early-stage PD patients and healthy individuals. The authors advocate for multimodal systems that combine different data types, such as speech, gait, and handwriting, to enhance diagnostic accuracy and reliability. They also address the interpretability of AI models through techniques like Grad-CAM and SHAP, which elucidate the decision-making processes of these systems, thereby fostering trust among clinicians. Overall, the findings underscore the potential of integrating diverse data sources and advanced machine learning techniques to improve early PD detection and clinical outcomes.
Limitations
The study acknowledges several limitations regarding the proposed multimodal framework. Although it shows enhanced robustness over unimodal methods when evaluated on retrospective benchmark datasets, its adaptability is constrained to feature-level fusion. Furthermore, the framework has yet to be validated in real-world dynamic clinical environments, which may affect its practical applicability and effectiveness in diverse clinical scenarios.
