DOI: https://doi.org/10.1007/s10462-024-10711-8
تاريخ النشر: 2024-02-26
المؤلف: Arti Anuragi وآخرون
الموضوع الرئيسي: تخطيط الدماغ وواجهات الدماغ-الكمبيوتر
نظرة عامة
يتناول هذا القسم من ورقة البحث تطبيق تقنيات الذكاء الاصطناعي (AI) للتحليل التلقائي لبيانات تخطيط الدماغ الكهربائي (EEG)، وهو أمر حاسم لتشخيص ومراقبة اضطرابات الدماغ. يبرز المؤلفون التحديات التي تطرحها مجموعات الميزات عالية الأبعاد المستمدة من إشارات EEG، وخاصة “لعنة الأبعاد”، التي يمكن أن تؤدي إلى الإفراط في التكيف وزيادة التعقيد في مهام التصنيف. لمعالجة هذه القضايا، تستعرض الورقة 23 تقنية لتقليل الأبعاد تهدف إلى تحويل ميزات EEG عالية الأبعاد إلى مساحة ميزات أكثر قابلية للإدارة، مما يعزز أداء التصنيف.
تقيّم المراجعة التجريبية فعالية هذه التقنيات لتقليل الأبعاد باستخدام مصنفات الذكاء الاصطناعي التقليدية، وبالتحديد آلة الدعم الناقل (SVM) وجار الأقرب (k-NN)، مع مقاييس الأداء بما في ذلك الدقة، الميزات المختارة، ومعدل تقليل الأبعاد (DRR). ومن الجدير بالذكر أن التقنية الهجينة القائمة على تحليل المكونات الرئيسية مع تضمين الجيران العشوائي الموزع (PCA+t-SNE) حققت دقة متوسطة قدرها 93.36% عند اقترانها بمصنف k-NN، متفوقة على المصنفات التي لم تستخدم تقنيات التقليل. تشير النتائج إلى أن PCA+t-SNE، إلى جانب تقنيات أخرى مثل F-ICA وUMAP، فعالة بشكل خاص في تصنيف مجموعات بيانات EEG عالية الأبعاد، مما يوفر رؤى قيمة للباحثين في اختيار طرق تقليل الأبعاد المناسبة لبيانات عالية الأبعاد مماثلة عبر مجالات مختلفة.
مقدمة
تناقش مقدمة الورقة الدور الحاسم لتخطيط الدماغ الكهربائي (EEG) في تشخيص وعلاج اضطرابات الدماغ، مع التأكيد على عدم كفاية الفحص البصري لتفسير بيانات EEG بدقة بسبب طبيعته المستغرقة للوقت وقابليته للخطأ البشري. لمعالجة هذه التحديات، يدعو المؤلفون إلى تطوير أنظمة التشخيص المدعومة بالكمبيوتر (CAD)، التي تقوم بأتمتة التقييم العصبي الفسيولوجي لإشارات EEG. تتكون هذه الأنظمة من أربعة مكونات رئيسية: معالجة البيانات، استخراج الميزات، تقليل الأبعاد، والتصنيف.
تحسن معالجة البيانات جودة بيانات EEG من خلال القضاء على الضوضاء وعدم التناسق، بينما يحدد استخراج الميزات الخصائص الرئيسية اللازمة لتمييز الأنشطة الدماغية. تبسط تقنيات تقليل الأبعاد، مثل اختيار الميزات وطرق التحويل مثل تحليل المكونات الرئيسية (PCA) وتحليل المكونات المستقلة (ICA)، التحليل وتحسن كفاءة النموذج. أخيرًا، يتضمن التصنيف تدريب نماذج التعلم الآلي لتسمية بيانات EEG بناءً على الميزات المتعلمة. يبرز المؤلفون أهمية طرق استخراج الميزات المختلفة، بما في ذلك تقنيات المجال الزمني، المجال الترددي، والزمن-الترددي، مشيرين إلى أن طرق الموجات المتقدمة كانت فعالة بشكل خاص في توليد متجهات ميزات عالية الأبعاد لتحليل اضطرابات الدماغ.
طرق
في قسم الطرق من الدراسة، يحدد المؤلفون مجموعات البيانات التجريبية المستخدمة لتقييم تقنيات الإسقاط، مع التركيز على ثلاث خصائص رئيسية: عدد الملاحظات (N)، الأبعاد (d)، ونسبة البعد الجوهري ($d_{intrinsic}$). يتم تصنيف عدد الملاحظات إلى ثلاث مجموعات: صغيرة (N < 1000)، متوسطة (1000 < N ≤ 4000)، وكبيرة (N > 4000)، مما يعكس أحجام مجموعات البيانات الشائعة في تقييمات طرق الإسقاط. يتم تصنيف الأبعاد بشكل مشابه إلى منخفضة (d < 100)، متوسطة (100 < d ≤ 500)، وعالية (d > 500). تُعرف نسبة البعد الجوهري، على أنها نسبة عدد المكونات الرئيسية التي تلتقط 95% من التباين إلى العدد الإجمالي للميزات، بأنها منخفضة ($d_{intrinsic} < 0.1$)، متوسطة ($0.1 < d_{intrinsic} ≤ 0.5$)، وعالية ($0.5 < d_{intrinsic} ≤ 1$)، مع القيم الأعلى تشير إلى تحديات في رسم البيانات. تستخدم الدراسة خمس مجموعات بيانات EEG المعالجة مسبقًا التي تتميز بمتجهات ميزات عالية الأبعاد تظهر مشكلات الأبعاد، كما هو ملاحظ في الأعمال السابقة (Anuragi 2023). يتم تقديم أوصاف مفصلة لهذه المجموعات، بما في ذلك أبعادها وعدد الملاحظات، في الجداول المرفقة، مما يسهل التحليل الشامل لإشارات EEG فيما يتعلق بالخصائص المحددة.
نتائج
تقدم نتائج هذه الدراسة التجريبية تحليلًا مقارنًا لـ 23 تقنية إسقاط تطبق على خمس ميزات لإشارات EEG عالية الأبعاد لتصنيف الفصام، والإدمان، والحالات البؤرية، والاكتئاب. باستخدام كل من Python وMATLAB، استخدمت الدراسة طرق معالجة مسبقة مختلفة لاستخراج الميزات، مع تقديم تفاصيل الأبعاد في الجدول 2. تم تقييم تقنيات الإسقاط، المفصلة في الجدول 3، باستخدام مصنفات k-NN وSVM مع نهج التحقق المتقاطع 10 مرات، مع تلخيص مقاييس الأداء في الجداول 7 و8. ومن الجدير بالذكر أن PCA مع t-SNE حقق دقة تصنيف متوسطة قدرها 93.36%، متفوقًا على متجه الميزات الكامل ومظهرًا فعالية تقليل الأبعاد.
كشفت التحليلات أن PCA+t-SNE، LLE، F-ICA، وISOMAP أدت أداءً جيدًا باستمرار عبر معظم مجموعات البيانات، بينما أظهرت NMDS-E، NMDS-M، LLP، وkernel PCA أداءً أقل. تشير النتائج إلى أن PCA+t-SNE فعالة بشكل خاص لمجموعات البيانات ذات الأبعاد الجوهري المنخفضة، مثل الفصام، والحالات البؤرية، والاكتئاب. توضح التصورات لأفضل أربع تقنيات تقليل، خاصة من مجموعة بيانات الاكتئاب، أن PCA+t-SNE أنتجت أكثر التضمينات معنى، مميزة بفعالية بين إشارات EEG الطبيعية والمكتئبة. بشكل عام، تؤكد النتائج فعالية تقنيات الإسقاط المحددة في تعزيز دقة التصنيف لتحليل إشارات EEG.
مناقشة
في هذا القسم، تناقش الورقة الدور الحاسم لتقليل الأبعاد في معالجة التحديات التي تطرحها البيانات عالية الأبعاد في نماذج التعلم الآلي المعتمدة على الذكاء الاصطناعي، وخاصة في سياق تصنيف إشارات EEG. يبرز المؤلفون “لعنة الأبعاد”، التي يمكن أن تؤدي إلى الإفراط في التكيف وانخفاض دقة التصنيف عندما تزداد أبعاد مساحة الميزات دون زيادة مقابلة في بيانات التدريب. للتخفيف من هذه القضايا، يتم التأكيد على تقنيات تقليل الأبعاد، وبالتحديد طرق تقليل الميزات، كأمر أساسي لتعزيز أداء النموذج من خلال تقليل التعدد الخطي وتصفيه الضوضاء في بيانات EEG.
تقدم المراجعة تحليلًا شاملاً لـ 23 تقنية إسقاط مختلفة، مصنفة إلى أنواع خطية وغير خطية، وتقيّم فعاليتها عبر مجموعات بيانات EEG المختلفة. تشير النتائج الرئيسية إلى أن طرق مثل PCA-t-SNE تتفوق على غيرها في تقليل الأبعاد مع الحفاظ على دقة التصنيف. تلخص الورقة أيضًا النتائج التجريبية من العديد من الدراسات، مما يوضح أن تقنيات تقليل الأبعاد المختلفة، بما في ذلك PCA، ICA، وt-SNE، قد تم تطبيقها بنجاح لتحسين نتائج التصنيف لمختلف اضطرابات الدماغ. يخلص المؤلفون إلى أن استكشاف هذه التقنيات بشكل أكبر ضروري لتحسين اختيار الميزات وتعزيز دقة نماذج التعلم الآلي في تحليل إشارات EEG.
DOI: https://doi.org/10.1007/s10462-024-10711-8
Publication Date: 2024-02-26
Author(s): Arti Anuragi et al.
Primary Topic: EEG and Brain-Computer Interfaces
Overview
This section of the research paper discusses the application of artificial intelligence (AI) techniques for the automatic analysis of electroencephalography (EEG) data, which is crucial for diagnosing and monitoring brain disorders. The authors highlight the challenges posed by high-dimensional feature sets derived from EEG signals, particularly the “curse of dimensionality,” which can lead to overfitting and increased complexity in classification tasks. To address these issues, the paper reviews 23 dimensionality reduction techniques aimed at transforming high-dimensional EEG features into a more manageable feature space, thereby enhancing classification performance.
The empirical review evaluates the effectiveness of these dimensionality reduction techniques using traditional AI classifiers, specifically support vector machine (SVM) and k-nearest neighbor (k-NN), with performance metrics including accuracy, selected features, and dimensionality reduction rate (DRR). Notably, the hybrid technique of principal component analysis-based t-distributed stochastic neighbor embedding (PCA+t-SNE) achieved an average accuracy of 93.36% when paired with the k-NN classifier, outperforming classifiers that did not utilize reduction techniques. The findings indicate that PCA+t-SNE, along with other techniques like F-ICA and UMAP, are particularly effective for classifying high-dimensional EEG datasets, providing valuable insights for researchers in selecting appropriate dimensionality reduction methods for similar high-dimensional data across various domains.
Introduction
The introduction of the paper discusses the critical role of electroencephalography (EEG) in diagnosing and treating brain disorders, emphasizing the inadequacy of visual inspection for accurate EEG data interpretation due to its time-consuming nature and susceptibility to human error. To address these challenges, the authors advocate for the development of computer-aided diagnosis (CAD) systems, which automate the neurophysiological assessment of EEG signals. These systems consist of four main components: data preprocessing, feature extraction, dimensionality reduction, and classification.
Data preprocessing enhances EEG data quality by eliminating noise and inconsistencies, while feature extraction identifies key characteristics necessary for distinguishing brain activities. Dimensionality reduction techniques, such as feature selection and transformation methods like principal component analysis (PCA) and independent component analysis (ICA), simplify the analysis and improve model efficiency. Finally, classification involves training machine learning models to label EEG data based on learned features. The authors highlight the importance of various feature extraction methods, including time-domain, frequency-domain, and time-frequency techniques, noting that advanced wavelet methods have been particularly effective in generating high-dimensional feature vectors for analyzing brain disorders.
Methods
In the Methods section of the study, the authors outline the experimental datasets used to evaluate projection techniques, focusing on three key characteristics: the number of observations (N), dimensionality (d), and intrinsic dimension ratio ($d_{intrinsic}$). The number of observations is categorized into three groups: Small (N < 1000), Medium (1000 < N ≤ 4000), and Large (N > 4000), reflecting common dataset sizes in projection method evaluations. Dimensionality is similarly classified into Low (d < 100), Medium (100 < d ≤ 500), and High (d > 500). The intrinsic dimension ratio, defined as the ratio of the number of principal components capturing 95% of the variance to the total number of features, is categorized as Low ($d_{intrinsic} < 0.1$), Medium ($0.1 < d_{intrinsic} ≤ 0.5$), and High ($0.5 < d_{intrinsic} ≤ 1$), with higher values indicating challenges in data mapping. The study utilizes five pre-processed EEG datasets characterized by high-dimensional feature vectors that exhibit dimensionality issues, as noted in previous work (Anuragi 2023). Detailed descriptions of these datasets, including their dimensionality and number of observations, are provided in accompanying tables, which facilitate a comprehensive analysis of EEG signals in relation to the defined characteristics.
Results
The results of this empirical study present a comparative analysis of 23 projection techniques applied to five high-dimensional EEG signal features for classifying schizophrenia, alcoholism, focal conditions, and depression. Utilizing both Python and MATLAB, the study employed various pre-processing methods to extract features, with dimensionality details provided in Table 2. The projection techniques, detailed in Table 3, were evaluated using k-NN and SVM classifiers with a 10-fold cross-validation approach, with performance metrics summarized in Tables 7 and 8. Notably, PCA combined with t-SNE achieved an average classification accuracy of 93.36%, outperforming the full feature vector and demonstrating effective dimensionality reduction.
The analysis revealed that PCA+t-SNE, LLE, F-ICA, and ISOMAP consistently performed well across most datasets, while NMDS-E, NMDS-M, LLP, and kernel PCA exhibited poorer performance. The results indicate that PCA+t-SNE is particularly effective for datasets with low intrinsic dimensionality, such as schizophrenia, focal, and depression. Visualizations of the top four reduction techniques, especially from the depression dataset, illustrated that PCA+t-SNE produced the most meaningful embeddings, effectively distinguishing between normal and depressed EEG signals. Overall, the findings underscore the efficacy of specific projection techniques in enhancing classification accuracy for EEG signal analysis.
Discussion
In this section, the paper discusses the critical role of dimension reduction in addressing the challenges posed by high-dimensional data in AI-based machine learning models, particularly in the context of EEG signal classification. The authors highlight the “curse of dimensionality,” which can lead to overfitting and decreased classification accuracy when the dimensionality of the feature space increases without a corresponding increase in training data. To mitigate these issues, dimension reduction techniques, specifically feature reduction methods, are emphasized as essential for enhancing model performance by reducing multicollinearity and filtering noise in EEG data.
The review presents an extensive analysis of 23 different projection techniques, categorizing them into linear and non-linear types, and evaluates their effectiveness across various EEG datasets. Key findings indicate that methods such as PCA-t-SNE outperform others in reducing dimensionality while maintaining classification accuracy. The paper also summarizes empirical results from numerous studies, demonstrating that various dimensionality reduction techniques, including PCA, ICA, and t-SNE, have been successfully applied to improve classification outcomes for different brain disorders. The authors conclude that further exploration of these techniques is necessary to optimize feature selection and enhance the accuracy of machine learning models in EEG signal analysis.
