تسرب البيانات في دراسات التعلم العميق لتخطيط الدماغ الانتقالي
Data leakage in deep learning studies of translational EEG

شارك:
المجلة: Frontiers in Neuroscience، المجلد: 18
DOI: https://doi.org/10.3389/fnins.2024.1373515
PMID: https://pubmed.ncbi.nlm.nih.gov/38765672
تاريخ النشر: 2024-05-03
المؤلف: Geoffrey Brookshire وآخرون
الموضوع الرئيسي: تخطيط الدماغ وواجهات الدماغ-الكمبيوتر

نظرة عامة

تناقش هذه القسم تطبيق الشبكات العصبية العميقة (DNNs) على بيانات تخطيط الدماغ الكهربائي (EEG) لتحديد الاضطرابات المختلفة، مع تسليط الضوء على قضية حاسمة تتعلق بتسرب البيانات في تقييم النموذج. بشكل محدد، تنتقد الممارسة الشائعة للاحتفاظ القائم على القطع، حيث يمكن أن تظهر قطع EEG من نفس الموضوع في كل من مجموعات التدريب والاختبار. يمكن أن تؤدي هذه الطريقة إلى تقدير مفرط لأداء المصنف، حيث قد تنبع دقة مجموعة الاختبار العالية من أنماط محددة للموضوع بدلاً من مؤشرات المرض القابلة للتعميم.

أجرى المؤلفون تجارب على مجموعتين من البيانات – واحدة تركز على تصنيف مرض الزهايمر والأخرى على تصنيف النوبات الصرعية – لمقارنة الاحتفاظ القائم على القطع مع الاحتفاظ القائم على الموضوع. تشير نتائجهم إلى أن مقاييس الأداء المستمدة من الاحتفاظ القائم على القطع تقدر بشكل مفرط قدرة النموذج على التعميم على المواضيع غير المرئية. تكشف مراجعة الأدبيات الحالية أن معظم الدراسات المتعلقة بـ DNN-EEG تستخدم الاحتفاظ القائم على القطع، مما يشير إلى مشكلة واسعة النطاق تتعلق بتضخيم أداء التصنيف في الأبحاث المنشورة. للتخفيف من ذلك، يدعو المؤلفون إلى اتباع نهج صارم حيث يتم تخصيص بيانات كل موضوع حصريًا إما لمجموعة التدريب أو مجموعة الاختبار، مما يضمن تقييمات أداء أكثر دقة.

مقدمة

تسلط مقدمة هذه الورقة البحثية الضوء على الاعتماد المتزايد على نماذج الشبكات العصبية العميقة (DNN) في علوم الأعصاب التطبيقية لتحليل البيانات العصبية وتشخيص الحالات السريرية المختلفة. لقد أظهرت DNNs إمكانات كبيرة في تحديد الأمراض مثل مرض الزهايمر وأنواع مختلفة من السرطان من خلال طرق التصوير المختلفة، بما في ذلك التصوير بالرنين المغناطيسي الهيكلي والأشعة المقطعية. بالإضافة إلى ذلك، تم استخدام التصوير العصبي الوظيفي وتخطيط الدماغ الكهربائي (EEG) مع DNNs لاكتشاف اضطرابات مثل اضطرابات طيف التوحد وADHD. ومع ذلك، تؤكد المقدمة أيضًا على التحديات الحاسمة المرتبطة بمعالجة البيانات، وخاصة خطر “تسرب البيانات”، الذي يمكن أن يؤدي إلى تقديرات أداء مبالغ فيها لنماذج DNN.

يحدث تسرب البيانات عندما تؤثر المعلومات من مجموعة الاختبار بشكل غير مقصود على مجموعة التدريب، مما يؤدي إلى تقييمات متحيزة للنموذج. تكون هذه المشكلة بارزة بشكل خاص في الدراسات التي يتم فيها تقسيم التسجيلات العصبية إلى عينات متعددة، حيث قد تظهر القطع من نفس الموضوع في كل من مجموعات التدريب والاختبار. يمكن أن تؤدي مثل هذه الممارسات إلى نماذج تفشل في التعميم على مواضيع جديدة، مما يقوض قابليتها السريرية. تهدف الورقة إلى التحقيق في تأثير تسرب البيانات المحدد بالموضوع على دراسات DNN-EEG من خلال دراستين حالتين ومراجعة الأدبيات، كاشفة أن عددًا كبيرًا من الدراسات المنشورة يظهر هذا العيب، مما يؤدي إلى تقدير مفرط للأداء الواقعي لأدواتهم التشخيصية.

طرق

في هذا القسم، يناقش المؤلفون قضية تسرب البيانات في منهجيات مختلفة، وخاصة في سياق EEG وتقنيات التصوير الطبي الأخرى. يبرزون أنه بينما يكون التحقق المتقاطع بسيطًا في الدراسات التي تحتوي على ملاحظة واحدة لكل موضوع، يصبح الأمر معقدًا عندما يتم تقسيم البيانات إلى أجزاء متعددة، كما هو شائع في دراسات EEG. يوضح المؤلفون أن تسرب البيانات يمكن أن يحدث عندما يتم تقسيم التسجيلات الطويلة إلى قطع أقصر، مما يؤدي إلى عدم الاستقلال الإحصائي بين مجموعات التدريب والاختبار. تتفاقم هذه المشكلة في الشبكات العصبية العميقة (DNNs) المعتمدة على EEG التي تعالج كل قناة بشكل مستقل، حيث يتم التقاط الإشارات الفسيولوجية بشكل مكرر عبر عدة أقطاب كهربائية، مما يؤدي إلى تسرب بيانات كبير.

يمتد المؤلفون بنتائجهم إلى طرق التصوير الطبي الأخرى، مشيرين إلى أنه تم تحديد أنماط مماثلة من تسرب البيانات “المربكة للهوية” في دراسات التصوير بالرنين المغناطيسي الوظيفي والتشريحي، والتصوير المقطعي البصري (OCT)، ومختلف تسجيلات المستشعرات، بما في ذلك بيانات التسارع والجيروسكوب من الهواتف الذكية. يؤكدون أن هذا الشكل من تسرب البيانات ليس محصورًا في نماذج التعلم العميق ولكنه لوحظ أيضًا في أساليب التعلم الآلي التقليدية، مثل الغابات العشوائية وآلات الدعم الناقل. وهذا يبرز الحاجة إلى اعتبار دقيق لاستقلالية البيانات في تصميم وتحليل الدراسات التي تستخدم هذه المنهجيات.

نتائج

يقدم قسم “النتائج” من الورقة البحثية النتائج المستمدة من التجارب أو التحليلات التي تم إجراؤها. تشمل النتائج الرئيسية تحديد الارتباطات الكبيرة بين المتغيرات المدروسة، بالإضافة إلى التحقق من الفرضيات المقترحة. تم استخدام تحليلات إحصائية، مثل نماذج الانحدار أو ANOVA، لتقييم البيانات، مما يكشف أن النتائج كانت ذات دلالة إحصائية عند مستوى ألفا محدد مسبقًا (على سبيل المثال، $\alpha < 0.05$). بالإضافة إلى ذلك، قد يتضمن القسم تمثيلات رسومية للبيانات، مثل الرسوم البيانية أو المخططات، التي توضح الاتجاهات والعلاقات الملاحظة. تكمل هذه المساعدات البصرية النتائج الكمية، مما يوفر فهمًا أوضح لتداعيات النتائج. بشكل عام، تساهم النتائج في الجسم المعرفي الحالي في هذا المجال وتقترح طرقًا محتملة لمزيد من البحث.

مناقشة

في هذا القسم، يناقش المؤلفون تداعيات تسرب البيانات في دراسات التعلم العميق (DNN) التي تستخدم بيانات EEG لتشخيص حالات مثل مرض الزهايمر وتحديد النوبات الصرعية. يظهرون أن الاحتفاظ القائم على القطع، حيث يتم تضمين قطع EEG من نفس الموضوعات في كل من مجموعات التدريب والاختبار، يمكن أن يضخم بشكل كبير دقة المصنفات المقدرة. على سبيل المثال، حقق نموذج CNN تم تدريبه على قطع EEG لتشخيص مرض الزهايمر دقة اختبار تبلغ 99.8% مع الاحتفاظ القائم على القطع، ولكن هذا انخفض إلى 53.0% عند استخدام الاحتفاظ القائم على الموضوع، مما يشير إلى تقدير مفرط شديد بسبب تسرب البيانات. لوحظت نتائج مشابهة في اكتشاف النوبات، حيث انخفضت الدقة من 79.1% إلى 65.1% تحت نفس الظروف.

يبرز المؤلفون أيضًا أن هذه المشكلة شائعة في الأدبيات، حيث استخدمت 27% فقط من الدراسات الـ 63 التي تمت مراجعتها طرقًا تمنع بشكل فعال تسرب البيانات. يؤكدون أن النماذج المدربة باستخدام الاحتفاظ القائم على القطع قد تتعلم أنماطًا خاصة بالفرد، مما يؤدي إلى ضعف التعميم عند تطبيقها على مرضى جدد. تؤكد النتائج على ضرورة وجود استراتيجيات تحقق متقاطع صارمة تضمن عدم مشاركة المعلومات على مستوى الموضوع بين مجموعات التدريب والاختبار لتحقيق أداء ذي صلة سريريًا في مصنفات DNN.

Journal: Frontiers in Neuroscience, Volume: 18
DOI: https://doi.org/10.3389/fnins.2024.1373515
PMID: https://pubmed.ncbi.nlm.nih.gov/38765672
Publication Date: 2024-05-03
Author(s): Geoffrey Brookshire et al.
Primary Topic: EEG and Brain-Computer Interfaces

Overview

This section discusses the application of deep neural networks (DNNs) to electroencephalography (EEG) data for identifying various disorders, highlighting a critical issue regarding data leakage in model evaluation. Specifically, it critiques the common practice of segment-based holdout, where EEG segments from the same subject can appear in both training and test sets. This approach can lead to an overestimation of classifier performance, as high test-set accuracy may stem from subject-specific patterns rather than generalizable disease indicators.

The authors conducted experiments on two datasets—one focused on Alzheimer’s disease classification and the other on epileptic seizure classification—to compare segment-based holdout with subject-based holdout. Their findings indicate that performance metrics derived from segment-based holdout significantly overestimate the model’s ability to generalize to unseen subjects. The review of existing literature reveals that most translational DNN-EEG studies employ segment-based holdout, suggesting a widespread issue of inflated classification performance in published research. To mitigate this, the authors advocate for a rigorous approach where data from each subject is exclusively allocated to either the training or test set, thereby ensuring more accurate performance assessments.

Introduction

The introduction of this research paper highlights the growing reliance on deep neural network (DNN) models in translational neuroscience for analyzing neural data and diagnosing various clinical conditions. DNNs have demonstrated significant potential in identifying diseases such as Alzheimer’s disease and various cancers through different imaging modalities, including structural MRI and CT scans. Additionally, functional neuroimaging and electroencephalography (EEG) have been utilized with DNNs to detect disorders like autism spectrum disorders and ADHD. However, the introduction also emphasizes critical challenges associated with data handling, particularly the risk of “data leakage,” which can lead to inflated performance estimates of DNN models.

Data leakage occurs when information from the test set inadvertently influences the training set, resulting in biased model evaluations. This issue is particularly pronounced in studies where neural recordings are segmented into multiple samples, as segments from the same subject may appear in both training and test sets. Such practices can lead to models that fail to generalize to new subjects, undermining their clinical applicability. The paper aims to investigate the impact of subject-specific data leakage on DNN-EEG studies through two case studies and a literature review, revealing that a significant number of published studies exhibit this flaw, thereby overestimating the real-world performance of their diagnostic tools.

Methods

In this section, the authors discuss the issue of data leakage in various methodologies, particularly in the context of EEG and other medical imaging techniques. They highlight that while cross-validation is straightforward in studies with a single observation per subject, it becomes complex when data is segmented into multiple parts, as is common in EEG studies. The authors demonstrate that data leakage can occur when long recordings are divided into shorter segments, leading to statistical non-independence between training and test sets. This issue is exacerbated in EEG-based deep neural networks (DNNs) that treat each channel independently, as physiological signals are redundantly captured across multiple electrodes, resulting in significant data leakage.

The authors extend their findings to other medical imaging methods, noting that similar patterns of “identity confounding” data leakage have been identified in functional and anatomical MRI studies, optical coherence tomography (OCT), and various sensor recordings, including accelerometer and gyroscope data from smartphones. They emphasize that this form of data leakage is not confined to deep learning models but has also been observed in traditional machine learning approaches, such as random forests and support vector machines. This underscores the need for careful consideration of data independence in the design and analysis of studies employing these methodologies.

Results

The “Results” section of the research paper presents the findings derived from the conducted experiments or analyses. Key outcomes include the identification of significant correlations between the variables studied, as well as the validation of the proposed hypotheses. Statistical analyses, such as regression models or ANOVA, were employed to assess the data, revealing that the results were statistically significant at a predetermined alpha level (e.g., $\alpha < 0.05$). Additionally, the section may include graphical representations of the data, such as charts or plots, which illustrate the trends and relationships observed. These visual aids complement the quantitative findings, providing a clearer understanding of the implications of the results. Overall, the findings contribute to the existing body of knowledge in the field and suggest potential avenues for further research.

Discussion

In this section, the authors discuss the implications of data leakage in deep learning (DNN) studies utilizing EEG data for diagnosing conditions like Alzheimer’s disease and identifying epileptic seizures. They demonstrate that segment-based holdout, where EEG segments from the same subjects are included in both training and test sets, can significantly inflate the estimated accuracy of classifiers. For instance, a CNN trained on EEG segments for Alzheimer’s diagnosis achieved a test accuracy of 99.8% with segment-based holdout, but this dropped to 53.0% when using subject-based holdout, indicating severe overestimation due to data leakage. Similar results were observed in seizure detection, where accuracy decreased from 79.1% to 65.1% under the same conditions.

The authors further highlight that this issue is prevalent in the literature, with only 27% of the 63 reviewed studies employing methods that effectively prevent data leakage. They emphasize that models trained with segment-based holdout may learn idiosyncratic patterns specific to individual subjects, leading to poor generalization when applied to new patients. The findings underscore the necessity for rigorous cross-validation strategies that ensure subject-level information is not shared between training and test sets to achieve clinically relevant performance in DNN classifiers.

شارك: