DOI: https://doi.org/10.1038/s41598-024-51985-w
PMID: https://pubmed.ncbi.nlm.nih.gov/38302557
تاريخ النشر: 2024-02-01
المؤلف: Abbas Saad Alatrany وآخرون
الموضوع الرئيسي: أبحاث الخرف وضعف الإدراك
نظرة عامة
تشخيص مرض الزهايمر (AD) في مراحله المبكرة معقد بسبب التغيرات الدقيقة في المؤشرات الحيوية التي غالبًا ما يتم تجاهلها. تستكشف هذه الدراسة تطبيق نماذج التعلم الآلي (ML)، وبشكل خاص آلات الدعم الناقل (SVM)، لتحديد الأفراد المعرضين لخطر مرض الزهايمر باستخدام مجموعة بيانات من المركز الوطني لتنسيق الزهايمر، والتي تشمل 169,408 سجلات و1024 ميزة. بعد تقليل مساحة الميزات، أظهرت نماذج SVM أداءً استثنائيًا، محققةً درجة F1 تبلغ 98.9% للتصنيف الثنائي (التمييز بين الإدراك الطبيعي (NC) وAD) و90.7% للتصنيف متعدد الفئات. بالإضافة إلى ذلك، توقعت النماذج تقدم مرض الزهايمر على مدى أربع سنوات، مع درجات F1 تبلغ 88% للمهام الثنائية و72.8% للمهام متعددة الفئات.
لمعالجة الحاجة الملحة لشرح النموذج، نفذت الدراسة طريقتين لاستخراج القواعد: تعدين قواعد الفئات ومجموعات القواعد المستقرة والقابلة للتفسير. أنتجت هذه الطرق قواعد مفهومة تساعد الخبراء في المجال على فهم العوامل الرئيسية التي تؤثر على تطور مرض الزهايمر. أكدت عملية التحقق من هذه القواعد من خلال نماذج SHAP وLIME على أهمية عوامل مثل الذاكرة، الحكم، التواصل، والتوجه في تقييم خطر مرض الزهايمر. كما أكدت النتائج على الدور المهم لأداة تقييم الخرف السريرية في توقع تقدم مرض الزهايمر.
الطرق
تستخدم البحث مزيجًا من اختيار الميزات وخوارزميات التعلم الآلي (ML) لتحديد الميزات المهمة المرتبطة بتقدم مرض الزهايمر (AD)، باستخدام مجموعة بيانات من المركز الوطني لتنسيق الزهايمر (NACC). تشمل المنهجية المعالجة المسبقة لإزالة القيم الشاذة والقيم المفقودة، تليها تقليل الأبعاد من خلال تحليل الارتباط وخوارزمية بوروتا. تهدف الدراسة إلى تصنيف وتوقع مراحل مرض الزهايمر مع استخراج رؤى قابلة للفهم من نماذج التعلم الآلي، مما يسهم في تحسين رعاية المرضى واستراتيجيات العلاج.
يتكون التصميم التجريبي من خمس تجارب رئيسية. التجربة 1 (EXP1) تقيم فعالية نماذج التعلم الآلي في تصنيف مراحل ضعف الإدراك باستخدام مجموعة الميزات الكاملة. التجربة 2 (EXP2) تقيم نفس النماذج على مجموعة ميزات مخفضة تم تحديدها بواسطة خوارزمية اختيار الميزات، مما يظهر فعاليتها في تصنيف مرض الزهايمر. التجربة 3 (EXP3) توسع هذا التحليل من خلال توقع الحالات الإدراكية بعد أربع سنوات من الأساس، واختبار القوة التنبؤية طويلة المدى للميزات المحددة. التجربة 4 (EXP4) تفحص قدرات التعميم للنماذج ذات الأداء العالي باستخدام مجموعة بيانات خارجية (ADNI)، بينما تستخدم التجربة 5 (EXP5) خوارزميات CAR وSIURS لاستخراج قواعد قابلة للقراءة البشرية تلتقط أنماط إدراكية مهمة. يتم التحقق من قوة الميزات المختارة من خلال مقارنة النتائج من CAR وSIURS مع نماذج SHAP وLIME.
النتائج
تظهر النتائج من التجارب (EXP1 إلى EXP5) فعالية مختلف مصنفات التعلم الآلي (ML) في التمييز بين الحالات الإدراكية، وبشكل خاص الإدراك الطبيعي (NC)، وضعف الإدراك الخفيف (MCI)، ومرض الزهايمر (AD). في EXP1، حقق خوارزم الغابة العشوائية (RF) أعلى دقة تبلغ 97.8% في تصنيف NC مقابل AD، بالإضافة إلى مقاييس أداء قوية مثل الدقة (97.2%)، الاسترجاع (98.1%)، ودرجة F1 (97.6%). تفوق نموذج RF باستمرار على المصنفات الأخرى عبر تصنيفات ثنائية مختلفة، بما في ذلك NC مقابل MCI وMCI مقابل AD، مما يدل على قوته وموثوقيته. بالمقابل، أظهرت مصنفات الجيران الأقرب (KNN) ونايف بايز (NB) أداءً أقل، خاصة في تصنيف حالات NC وMCI.
في EXP2، تم تقييم المصنفات باستخدام مجموعة ميزات مخفضة، حيث تصدرت RF مرة أخرى بدقة تبلغ 97.5% لـ NC مقابل AD. أظهرت كل من نماذج RF وآلات الدعم الناقل (SVM) أداءً عاليًا عبر جميع مهام التصنيف، بينما تأخرت نموذج NB. ركزت EXP3 على توقع الحالات الإدراكية بعد أربع سنوات من الزيارة الأولية، حيث حافظت RF على تفوقها بدقة تبلغ 96.4% لـ NC مقابل AD. أخيرًا، قامت EXP4 بالتحقق من المصنفات باستخدام مجموعة بيانات ADNI، مؤكدةً عمومية نماذج RF وSVM، التي حققت دقة تبلغ 98.3% و99%، على التوالي، لتصنيف NC مقابل AD. استخدمت EXP5 خوارزميات قواعد الارتباط الفئوية (CAR) ومجموعة القواعد المستقرة والقابلة للتفسير (SIRUS) لاستخراج قواعد قابلة للقراءة البشرية، كاشفةً عن متغيرات رئيسية مثل الذاكرة، الحكم، والتواصل التي تؤثر بشكل كبير على تصنيف مرض الزهايمر. بشكل عام، تؤكد النتائج على إمكانية استخدام RF وSVM كأدوات موثوقة لتصنيف الحالة الإدراكية، مدعومةً بقواعد قابلة للتفسير تعزز فهم أنماط البيانات الأساسية.
المناقشة
تسلط قسم المناقشة في ورقة البحث الضوء على فعالية نماذج التعلم الآلي (ML) المختلفة—وبشكل خاص الغابة العشوائية (RF)، وآلات الدعم الناقل (SVM)، ونايف بايز (NB)، والجيران الأقرب (KNN)—في تصنيف مرضى الزهايمر (AD) من أولئك الذين لديهم إدراك طبيعي (NC) أو ضعف إدراك خفيف (MCI). تشير النتائج إلى أن RF وSVM حققت باستمرار أعلى دقة، ودقة، واسترجاع عبر جميع مهام التصنيف، مما يظهر قدرتهما الفائقة في التمييز بين NC وMCI وAD. بالمقابل، بينما حقق NB وKNN أيضًا أداءً جيدًا، إلا أنهما عمومًا تأخرا عن RF وSVM، مما يشير إلى أن النماذج السابقة أكثر ملاءمة للتشخيص الدقيق لمرض الزهايمر.
تؤكد الدراسة أيضًا على قوة SVM، التي لم تتفوق فقط في مهام التصنيف ولكن أيضًا حافظت على أداء عالٍ في المهام التنبؤية عبر مجموعات بيانات مختلفة، بما في ذلك التحقق الخارجي باستخدام مجموعة بيانات مبادرة تصوير مرض الزهايمر (ADNI). أثبتت عملية اختيار الميزات، التي قللت مجموعة البيانات من 64 إلى 21 ميزة، فعاليتها في تعزيز أداء النموذج مع الحفاظ على القابلية للتفسير. استخدمت البحث أيضًا طرق استخراج القواعد (CARs وSIRUS) لتحديد الميزات المهمة التي تؤثر على قرارات النموذج، كاشفةً أن الذاكرة، الحكم، والتوجه هي عوامل حاسمة في تقييم خطر مرض الزهايمر. يعزز توافق أهمية الميزات عبر نماذج مختلفة موثوقية النتائج، مما يشير إلى أن هذه المصنفات ML يمكن أن تتنبأ بشكل موثوق بتغيرات الصحة الإدراكية، وبالتالي توفير رؤى قيمة للتطبيقات السريرية في تشخيص وإدارة مرض الزهايمر.
القيود
في هذا القسم، يعترف المؤلفون بعدة قيود لدراستهم. يشيرون إلى أن تحليل ارتباط الميزات تم على قيم مستمرة مفصولة، مما قد يؤدي إلى فقدان المعلومات ويؤثر على دقة تقييمات الارتباط. بالإضافة إلى ذلك، قد يؤدي الاعتماد على تسمية تشخيصية مستمدة سريريًا إلى تقييم غير مكتمل لأداء النموذج، حيث قد لا يأخذ في الاعتبار جميع المتغيرات ذات الصلة.
مع النظر إلى المستقبل، يقترح المؤلفون اتجاهات بحثية مستقبلية تشمل إجراء تحليلات طولية لفحص تطور درجات تقييم الخرف السريرية (CDR) بمرور الوقت. يهدف هذا النهج إلى تعزيز الكشف المبكر عن مرض الزهايمر (AD) وتسهيل المراقبة المستمرة. كما يخططون لتوسيع منهجيتهم لتشمل دراسات أكبر تركز على توقع مراحل المرض المبكرة، وخاصة الانتقال من حالات غير خرفية إلى حالات خرفية. علاوة على ذلك، يعتزم المؤلفون استكشاف استخدام نماذج التجميع التي تدمج عدة خوارزميات قابلة للتفسير، مما قد يحسن من قوة وموثوقية تفسيرات نماذجهم.
DOI: https://doi.org/10.1038/s41598-024-51985-w
PMID: https://pubmed.ncbi.nlm.nih.gov/38302557
Publication Date: 2024-02-01
Author(s): Abbas Saad Alatrany et al.
Primary Topic: Dementia and Cognitive Impairment Research
Overview
The early diagnosis of Alzheimer’s disease (AD) is complicated by subtle biomarker changes that are often missed. This study explores the application of machine learning (ML) models, specifically support vector machines (SVM), to identify individuals at risk for AD using a dataset from the National Alzheimer’s Coordinating Center, which includes 169,408 records and 1024 features. After feature space reduction, the SVM models demonstrated exceptional performance, achieving an F1 score of 98.9% for binary classification (distinguishing between normal cognition (NC) and AD) and 90.7% for multiclass classification. Additionally, the models effectively predicted AD progression over four years, with F1 scores of 88% for binary tasks and 72.8% for multiclass tasks.
To address the critical need for model explainability, the study implemented two rule-extraction methods: class rule mining and stable and interpretable rule sets. These methods produced comprehensible rules that aid domain experts in understanding the key factors influencing AD development. Validation of these rules through SHAP and LIME models highlighted the importance of factors such as MEMORY, JUDGMENT, COMMUN, and ORIENT in assessing AD risk. The findings also emphasized the Clinical Dementia Rating tool’s significant role in predicting AD progression.
Methods
The research employs a combination of feature selection and machine learning (ML) algorithms to identify significant features associated with Alzheimer’s Disease (AD) progression, utilizing a dataset from the National Alzheimer’s Coordinating Center (NACC). The methodology includes pre-processing to eliminate outliers and missing values, followed by dimensionality reduction through correlation analysis and the Boruta algorithm. The study aims to classify and predict AD stages while deriving human-understandable insights from the ML models, ultimately contributing to improved patient care and treatment strategies.
The experimental design consists of five key experiments. Experiment 1 (EXP1) assesses the efficacy of ML models in classifying cognitive impairment stages using the full feature set. Experiment 2 (EXP2) evaluates the same models on a reduced feature set identified by the feature selection algorithm, demonstrating its effectiveness in AD classification. Experiment 3 (EXP3) extends this analysis by predicting cognitive states four years post-baseline, testing the long-term predictive power of the identified features. Experiment 4 (EXP4) examines the generalization capabilities of the top-performing models using an external dataset (ADNI), while Experiment 5 (EXP5) employs the CAR and SIURS algorithms to extract human-readable rules that capture significant cognitive patterns. The robustness of the selected features is further validated by comparing results from CAR and SIURS with SHAP and LIME models.
Results
The results from the experiments (EXP1 to EXP5) demonstrate the effectiveness of various machine learning (ML) classifiers in distinguishing between cognitive states, specifically normal cognition (NC), mild cognitive impairment (MCI), and Alzheimer’s disease (AD). In EXP1, the Random Forest (RF) algorithm achieved the highest accuracy of 97.8% for classifying NC versus AD, along with strong performance metrics such as precision (97.2%), recall (98.1%), and F1 score (97.6%). The RF model consistently outperformed other classifiers across different pairwise classifications, including NC vs MCI and MCI vs AD, indicating its robustness and reliability. In contrast, the K-Nearest Neighbors (KNN) and Naive Bayes (NB) classifiers exhibited poorer performance, particularly in classifying NC and MCI cases.
In EXP2, the classifiers were evaluated using a reduced feature set, with RF again leading with an accuracy of 97.5% for NC vs AD. Both RF and Support Vector Machine (SVM) models demonstrated high performance across all classification tasks, while the NB model lagged behind. EXP3 focused on predicting cognitive states four years post-initial visit, where RF maintained its superiority with an accuracy of 96.4% for NC vs AD. Finally, EXP4 validated the classifiers using the ADNI dataset, confirming the generalizability of the RF and SVM models, which achieved accuracies of 98.3% and 99%, respectively, for NC vs AD classification. EXP5 employed the Class Association Rules (CAR) and Stable and Interpretable Rule Set (SIRUS) algorithms to extract human-readable rules, revealing key variables such as MEMORY, JUDGMENT, and COMMUN that significantly influence AD classification. Overall, the findings underscore the potential of RF and SVM as reliable tools for cognitive state classification, supported by interpretable rules that enhance understanding of the underlying data patterns.
Discussion
The discussion section of the research paper highlights the effectiveness of various machine learning (ML) models—specifically Random Forest (RF), Support Vector Machine (SVM), Naive Bayes (NB), and K-Nearest Neighbors (KNN)—in classifying Alzheimer’s Disease (AD) patients from those with Normal Cognition (NC) or Mild Cognitive Impairment (MCI). The results indicate that RF and SVM consistently achieved the highest accuracy, precision, and recall across all classification tasks, demonstrating their superior capability in distinguishing between NC, MCI, and AD. In contrast, while NB and KNN also performed well, they generally lagged behind RF and SVM, suggesting that the former models are more suitable for accurate AD diagnosis.
The study further emphasizes the robustness of SVM, which not only excelled in classification tasks but also maintained high performance in predictive tasks across different datasets, including external validation with the Alzheimer’s Disease Neuroimaging Initiative (ADNI) dataset. The feature selection process, which reduced the dataset from 64 to 21 features, proved effective in enhancing model performance while maintaining interpretability. The research also employed rule extraction methods (CARs and SIRUS) to identify significant features influencing model decisions, revealing that MEMORY, JUDGMENT, and ORIENT are critical in assessing AD risk. The alignment of feature importance across different models reinforces the reliability of the findings, suggesting that these ML classifiers can reliably forecast cognitive health changes, thus providing valuable insights for clinical applications in AD diagnosis and management.
Limitations
In this section, the authors acknowledge several limitations of their study. They note that the feature correlation analysis was performed on discretized continuous values, which may lead to information loss and impact the accuracy of the correlation assessments. Additionally, the reliance on a clinically derived diagnosis label could result in an incomplete evaluation of the model’s performance, as it may not account for all relevant variables.
Looking ahead, the authors propose future research directions that include conducting longitudinal analyses to examine the evolution of Clinical Dementia Rating (CDR) scores over time. This approach aims to enhance early detection of Alzheimer’s disease (AD) and facilitate continuous monitoring. They also plan to broaden their methodology to encompass larger studies focused on predicting early disease stages, particularly the transition from non-demented to demented states. Furthermore, the authors intend to explore the use of ensemble models that integrate multiple explainable algorithms, which could improve the robustness and reliability of their model explanations.
