التعلم الآلي القائم على الشبكات والتكامل متعدد الأوميات للتشخيص المبكر لمرض باركنسون
Integrative multi-omics and network-based machine learning for early diagnosis of Parkinson’s disease

شارك:
المجلة: PLoS ONE، المجلد: 21، العدد: 1
DOI: https://doi.org/10.1371/journal.pone.0329980
PMID: https://pubmed.ncbi.nlm.nih.gov/41493977
تاريخ النشر: 2026-01-06
المؤلف: Wei Liu وآخرون
الموضوع الرئيسي: آليات وعلاجات مرض باركنسون

نظرة عامة

تبحث الدراسة في دمج التعلم الآلي مع التحليلات متعددة الأوميات وتحليلات الشبكات لتحسين تشخيص مرض باركنسون (PD)، الذي يتعقد بسبب تعقيداته البيولوجية. كانت الدراسة تهدف إلى تصنيف مرضى PD في مراحل مبكرة والأشخاص الأصحاء (HCs) من خلال تحليل بيانات من 305 مشارك في مجموعة PPMI، باستخدام ميثيل الحمض النووي، وتعبير الجينات، وبيانات البروتينات. تم إجراء اختيار الميزات باستخدام تحليل التمييز الجزئي الأقل تفرعًا (sPLS-DA)، تلاه الدمج من خلال DIABLO. استخدمت الدراسة مصنفات XGBoost، محققة أداءً مثاليًا مع نموذج متعدد الأوميات الذي حقق منطقة تحت المنحنى (AUC) تبلغ 0.72 ودقة 0.74 في مجموعة الاختبار.

كشفت النتائج عن 56 موقع CpG مهم، و61 جينًا، و70 بروتينًا من خلال دمج DIABLO، إلى جانب تحديد 59 جينًا تنظيميًا رئيسيًا عبر تحليل الشبكة. بينما تفوق نموذج متعدد الأوميات على الآخرين في مجموعة الاختبار، أظهر نموذج المنظمين الطوبولوجيين أداءً متفوقًا في مجموعة التحقق مع AUC تبلغ 0.57 ودقة 0.62. بشكل عام، تؤكد الدراسة على إمكانيات دمج التعلم الآلي مع الأوميات التكاملي وتحليل الشبكات لتحديد العلامات الحيوية بشكل فعال وتصنيف PD، مما يشير إلى آفاق واعدة لتطبيقات التشخيص السريري واستراتيجيات العلاج المستقبلية.

مقدمة

يعتبر مرض باركنسون (PD) ثاني أكثر الاضطرابات التنكسية العصبية انتشارًا بين كبار السن، حيث يؤثر على ما يصل إلى 4% من الأفراد بحلول سن 80. يتميز بأعراض حركية، مثل بطء الحركة والصلابة، وميزات غير حركية مثل ضعف الإدراك، مما يؤثر بشكل كبير على جودة الحياة. على الرغم من الأبحاث الواسعة التي تكشف عن جوانب جزيئية ومرضية مختلفة لـ PD، إلا أنه لا يزال غير قابل للعلاج، حيث توفر العلاجات الحالية فقط تخفيفًا للأعراض. تعتبر التحديات الرئيسية في إدارة PD هي تشخيصه في مراحل متأخرة، والذي يحدث غالبًا بعد فقدان كبير للخلايا العصبية الدوبامينية، مما يبرز الحاجة الملحة لعلامات حيوية حساسة ودقيقة للكشف المبكر.

تتطلب هتروجينية PD، المتأثرة بالعوامل الوراثية، والوراثية، والتمثيل الغذائي، والالتهابية، دمج مؤشرات جزيئية متعددة لتحديد العلامات الحيوية بشكل فعال. لقد ظهرت تقنيات “الأوميات” عالية الإنتاجية – مثل الجينوميات، والتعبير الجيني، والبروتيوميات، والتمثيل الغذائي – كأدوات أساسية للتوصيف الشامل للتغيرات الجزيئية طوال تقدم المرض. تعزز تقنيات التعلم الآلي (ML) هذه العملية من خلال تحليل البيانات الطبية الحيوية عالية الأبعاد، مما يحسن دقة التشخيص وتوقع الأنماط الفرعية. في هذه الدراسة، استخدمنا ML لدمج مصادر البيانات البيولوجية المتنوعة، بما في ذلك مصفوفات ميثيل الحمض النووي، وبيانات التعبير الجيني، وملفات البروتيوميات، لتصنيف مرضى PD المبكرين من الأشخاص الأصحاء. لم يقتصر نهجنا على تحديد علامات حيوية مرشحة جديدة فحسب، بل أوضح أيضًا الروابط بين هذه العلامات عبر طبقات الأوميات المختلفة، مما يمهد الطريق للتقدم في التشخيصات السريرية والطب الدقيق.

طرق

في هذه الدراسة، شمل معالجة البيانات استخدام مصفوفات ميثيل الحمض النووي، وملفات التعبير الجيني، وبيانات البروتيوميات المستمدة من مبادرة علامات تقدم مرض باركنسون (PPMI) اعتبارًا من أكتوبر 2024. كانت المجموعة تتكون من مرضى باركنسون في مراحل مبكرة (PD) الذين لم يتلقوا علاجات دوبامينية ولم يحملوا طفرات في جينات LRRK2، وGBA، أو SNCA، بينما كانت مجموعة التحكم الصحية (HCs) متطابقة في العمر والجنس وكانت طبيعية عصبيًا وخالية من الأدوية التي تؤثر على النظام الدوباميني. تم تقسيم العينات إلى مجموعات تدريب (80%) واختبار (20%)، مع إجراء تحليلات بيولوجية معلوماتية أحادية الأوميات ومتعددة الأوميات على مجموعة التدريب.

تم تطبيق التعلم الآلي (ML) على مجموعة البيانات الكاملة، لكن تدريب النموذج استخدم فقط مجموعة التدريب، مع الاحتفاظ بمجموعة البيانات الاختبارية للتحقق الداخلي. لتقييم قابلية تعميم النموذج، تم دمج مجموعة بيانات تحقق خارجية (GSE165083) من الأرشيف الجيني (GEO)، والتي تضمنت ملفات ميثيل الحمض النووي وملفات التعبير من مرضى PD ومجموعة التحكم المتطابقة. تم تحليل العينات التي تحتوي على بيانات متعددة الأوميات كاملة فقط لتقليل التحيز في التحليلات اللاحقة، خاصة في تقدير المكونات الكامنة.

نتائج

يقدم قسم “النتائج” من ورقة البحث النتائج المستمدة من التجارب أو التحليلات التي تم إجراؤها. عادةً ما يتضمن بيانات كمية، وتحليلات إحصائية، وتمثيلات بصرية مثل الرسوم البيانية أو الجداول التي توضح نتائج الدراسة. غالبًا ما يتم تسليط الضوء على النتائج الرئيسية، مما يظهر الاتجاهات أو الأنماط المهمة التي تدعم الفرضيات أو أسئلة البحث المطروحة سابقًا في الورقة.

بالإضافة إلى تقديم البيانات الخام، قد يناقش هذا القسم أيضًا تداعيات النتائج، مقارنتها بالدراسات السابقة والأطر النظرية. عادةً ما يتم تأطير النتائج ضمن سياق أهداف البحث، مع التأكيد على كيفية مساهمتها في المعرفة الحالية في هذا المجال. بشكل عام، يخدم هذا القسم في التحقق من منهجية البحث وتوفير أساس للمناقشات والاستنتاجات اللاحقة.

مناقشة

في البحث الذي تم مناقشته، تم معالجة وتحليل بيانات ميثيل الحمض النووي، وتعبير الجينات، وبيانات البروتيوميات بدقة للتحقيق في أدوارها في مرض باركنسون (PD) باستخدام مجموعة PPMI وقاعدة بيانات GEO. خضعت بيانات ميثيل الحمض النووي للتطبيع عبر طريقة التطبيع الوظيفي، مع التركيز على أعلى 10,000 جين الأكثر تباينًا بعد استبعاد المتنبئين ذوي التباين المنخفض. تم معالجة بيانات التعبير الجيني باستخدام تطبيع DESeq2، مما أدى إلى اختيار أعلى 5,000 جين متغير. تم تحليل بيانات البروتيوميات من السائل الدماغي الشوكي باستخدام منصة SOMAscan، مع وجود تدابير صارمة لمراقبة الجودة. استخدمت التحليلات الأحادية الأوميات تحليل التمييز الجزئي الأقل تفرعًا (sPLS-DA) للتصنيف واختيار الميزات، بينما استخدمت التحليل التكاملي خوارزمية DIABLO لتحديد المكونات الكامنة المرتبطة عبر طبقات الأوميات.

شمل تحسين معلمات الدراسة لـ sPLS-DA تقييمًا منهجيًا لعدد المكونات وعدد الميزات، مع تقييم أداء النموذج من خلال التحقق المتقاطع المقسم. كشفت تحليل الإثراء الوظيفي عن عمليات بيولوجية مهمة تتعلق بالاستجابات المناعية والالتهابية عبر جميع المكونات، مع تحديد مسارات محددة لكل مكون. أظهر نموذج الأوميات المتكامل أداءً متفوقًا مقارنةً بنماذج الأوميات الأحادية، مع قوة تمييز قوية بين مرضى PD والأشخاص الأصحاء. سلط تحليل الشبكة الطوبولوجي الضوء على الجينات التنظيمية الرئيسية وأدوارها في المسارات المرتبطة بالتهاب الأعصاب والإجهاد التأكسدي، مما يبرز أهميتها المحتملة في مسببات PD. حققت نماذج التعلم الآلي، وخاصة نموذج توقيع الأوميات المتعددة، دقة تصنيف ملحوظة، مما يشير إلى إمكانيات الأساليب المتكاملة للأوميات في فهم الأمراض المعقدة مثل PD.

Journal: PLoS ONE, Volume: 21, Issue: 1
DOI: https://doi.org/10.1371/journal.pone.0329980
PMID: https://pubmed.ncbi.nlm.nih.gov/41493977
Publication Date: 2026-01-06
Author(s): Wei Liu et al.
Primary Topic: Parkinson\'s Disease Mechanisms and Treatments

Overview

The research investigates the integration of machine learning with multi-omics and network topological analyses to improve the diagnosis of Parkinson’s Disease (PD), which is complicated by its biological intricacies. The study aimed to classify early-stage PD patients and healthy controls (HCs) by analyzing data from 305 participants in the PPMI cohort, utilizing DNA methylation, gene expression, and proteomic data. Feature selection was performed using sparse Partial Least Squares Discriminant Analysis (sPLS-DA), followed by integration through DIABLO. The study employed XGBoost classifiers, achieving optimal performance with a multi-omics model that yielded an area under the curve (AUC) of 0.72 and accuracy of 0.74 in the test set.

The findings revealed 56 significant CpG sites, 61 genes, and 70 proteins through DIABLO integration, alongside the identification of 59 key regulatory genes via network topology analysis. While the multi-omics model outperformed others in the test set, the topological regulators model showed superior performance in the validation set with an AUC of 0.57 and accuracy of 0.62. Overall, the study underscores the potential of combining machine learning with integrative multi-omics and network analysis for effective biomarker identification and PD classification, suggesting promising avenues for clinical diagnostic applications and future therapeutic strategies.

Introduction

Parkinson’s disease (PD) is the second most prevalent neurodegenerative disorder in the elderly, affecting up to 4% of individuals by age 80. Characterized by both motor symptoms, such as bradykinesia and rigidity, and non-motor features like cognitive dysfunction, PD significantly impairs quality of life. Despite extensive research revealing various molecular and pathological aspects of PD, it remains incurable, with current treatments only providing symptomatic relief. A critical challenge in managing PD is its late-stage diagnosis, which often occurs after considerable loss of dopaminergic neurons, underscoring the urgent need for sensitive and specific biomarkers for early detection.

The heterogeneity of PD, influenced by genetic, epigenetic, metabolic, and inflammatory factors, necessitates the integration of multiple molecular indicators for effective biomarker identification. High-throughput “omics” technologies—such as genomics, transcriptomics, proteomics, and metabolomics—have emerged as essential tools for comprehensive profiling of molecular changes throughout disease progression. Machine learning (ML) techniques further enhance this process by analyzing high-dimensional biomedical data, improving diagnostic accuracy and subtype prediction. In this study, we utilized ML to integrate diverse biological data sources, including DNA methylation arrays, gene expression data, and proteomic profiles, to classify early PD patients from healthy controls. Our approach not only identified novel candidate biomarkers but also elucidated the interconnections among these markers across various omics layers, paving the way for advancements in clinical diagnostics and precision medicine.

Methods

In this study, data preprocessing involved the use of DNA methylation arrays, gene expression profiles, and proteomic data sourced from the Parkinson’s Progression Markers Initiative (PPMI) as of October 2024. The cohort consisted of early-stage Parkinson’s disease (PD) patients who had not received dopaminergic treatments and did not possess mutations in the LRRK2, GBA, or SNCA genes, while age- and sex-matched healthy controls (HCs) were neurologically normal and free from medications affecting the dopaminergic system. Samples were stratified into training (80%) and testing (20%) cohorts, with single-omics and multi-omics bioinformatic analyses conducted on the training set.

Machine learning (ML) was applied to the full dataset, but model training utilized only the training cohort, reserving the testing dataset for internal validation. To evaluate the model’s generalizability, an external validation dataset (GSE165083) from the Gene Expression Omnibus (GEO) was incorporated, which included DNA methylation and expression profiles from PD patients and matched controls. Only samples with complete multi-omics data were analyzed to reduce bias in subsequent analyses, particularly in latent component estimation.

Results

The “Results” section of the research paper presents the findings derived from the conducted experiments or analyses. It typically includes quantitative data, statistical analyses, and visual representations such as graphs or tables that illustrate the outcomes of the study. Key results are often highlighted, demonstrating significant trends or patterns that support the hypotheses or research questions posed earlier in the paper.

In addition to presenting raw data, this section may also discuss the implications of the findings, comparing them to previous studies and theoretical frameworks. The results are usually framed within the context of the research objectives, emphasizing how they contribute to the existing body of knowledge in the field. Overall, this section serves to validate the research methodology and provide a foundation for subsequent discussions and conclusions.

Discussion

In the discussed research, DNA methylation, gene expression, and proteomic data were meticulously processed and analyzed to investigate their roles in Parkinson’s disease (PD) using the PPMI cohort and GEO database. The DNA methylation data underwent normalization via the Functional Normalization method, focusing on the top 10,000 most variable genes after excluding low-variance predictors. Gene expression data were processed with DESeq2 normalization, leading to the selection of the top 5,000 variable genes. Proteomic data from cerebrospinal fluid were analyzed using the SOMAscan platform, with rigorous quality control measures in place. Single-omics analyses employed sparse partial least squares discriminant analysis (sPLS-DA) for classification and feature selection, while integrative analysis utilized the DIABLO algorithm to identify correlated latent components across the omics layers.

The study’s parameter optimization for sPLS-DA involved systematic evaluation of component numbers and feature counts, with model performance assessed through stratified cross-validation. Functional enrichment analysis revealed significant biological processes related to immune and inflammatory responses across all components, with specific pathways identified for each component. The integrative multi-omics model demonstrated superior performance compared to single-omics models, with a robust discriminatory power between PD patients and healthy controls. Network topological analysis highlighted key regulatory genes and their roles in pathways associated with neuroinflammation and oxidative stress, underscoring their potential relevance in PD pathogenesis. Machine learning models, particularly the multi-omics signature model, achieved notable classification accuracy, indicating the potential of integrated omics approaches in understanding complex diseases like PD.

شارك: