DOI: https://doi.org/10.1038/s41598-025-08362-y
PMID: https://pubmed.ncbi.nlm.nih.gov/40634379
تاريخ النشر: 2025-07-09
المؤلف: Sana Latif وآخرون
الموضوع الرئيسي: كشف وتصنيف أورام الدماغ
نظرة عامة
تتناول هذه الدراسة الحاجة الملحة للكشف المبكر عن مرض الزهايمر (AD)، وهو اضطراب عصبي تنكسي تدريجي يؤثر بشكل كبير على الوظائف الإدراكية، وخاصة فهم اللغة. الطرق التشخيصية الحالية محدودة، وتقترح الدراسة نهجًا جديدًا يستخدم نموذج تعلم عميق ذو فرعين يجمع بين مشفر BERT وشبكة عصبية تلافيفية متكررة (RCNN). يقوم مشفر BERT بمعالجة البيانات النصية من الملاحظات السريرية، محولًا إياها إلى متجهات ثابتة الطول، والتي يتم تحليلها بعد ذلك من خلال طبقات تلافيفية وLSTM للتنبؤ بوجود مرض الزهايمر. كما تعالج RCNN نفس بيانات النص، ويتم دمج المخرجات من كلا الفرعين باستخدام التعلم الجماعي لتعزيز دقة التصنيف.
تم تقييم النموذج المقترح على مجموعة Cookie Theft من مجموعة DementiaBank Pitt، محققًا دقة مثيرة للإعجاب تبلغ 94.98%، وF1-score قدره 0.9523، وAUC قدره 0.93. تشير هذه النتائج إلى أن النموذج يتفوق على التقنيات الحالية الرائدة في تشخيص مرض الزهايمر المبكر بناءً على تحليل النص. تبرز هذه الدراسة إمكانيات الذكاء الاصطناعي، وخاصة خوارزميات التعلم العميق، في تحسين دقة التشخيص وتسهيل التدخل في الوقت المناسب لمرض الزهايمر.
مقدمة
في هذا القسم، يقدم المؤلفون ثلاث ابتكارات مهمة في نهجهم للكشف عن مرض الزهايمر من النصوص السريرية، بناءً على الأبحاث الحالية التي استخدمت تجميعات على مستوى القرار من BERT وCNN لتصنيف النص.
أولاً، يقدمون بنية هجينة تدمج فروع BERT وCNN وLSTM، مما يمكّن النموذج من التقاط كل من الاعتماديات بعيدة المدى والأنماط التسلسلية المحلية في اللغة التي قد تشير إلى تدهور إدراكي. ثانيًا، يقومون بتنفيذ تجميع على مستوى الميزات مع وزن متكيف، حيث يتم دمج تمثيلات بُعد 128 من الطبقة قبل الأخيرة لكل فرع ومعالجتها من خلال شبكة عصبية متعددة الطبقات (MLP) ضحلة. تتعلم هذه الشبكة العصبية أوزانًا ديناميكية لكل ميزة، مما يسمح للنموذج بإعطاء الأولوية للفروع الأكثر معلوماتية بناءً على الإدخال المحدد. أخيرًا، يقترح المؤلفون خط أنابيب معالجة مسبقة مدرك سريريًا يتضمن تطبيع المفاهيم السريرية، وتوضيح معنى الكلمات، وتصحيح عدم سلاسة الكلام، مما يضمن الحفاظ بدقة على العلامات اللغوية ذات الصلة بمرض الزهايمر وإبرازها.
طرق البحث
تشمل المنهجية المقترحة في هذه الدراسة نموذج شبكة عصبية ذو فرعين مصمم لتصنيف مرض الزهايمر (AD) بناءً على بيانات نصية. يستخدم الفرع الأول مشفر BERT لتحويل النص المدخل إلى متجه ثابت الطول، والذي تتم معالجته بعد ذلك من خلال طبقة تلافيفية، وطبقة LSTM، وطبقة متصلة بالكامل للتنبؤ بوجود مرض الزهايمر. يستخدم الفرع الثاني شبكة عصبية تلافيفية متكررة (RCNN) تعالج أيضًا بيانات النص لإنتاج مخرجات التصنيف. يتم دمج هذين الفرعين باستخدام نهج التعلم الجماعي، مما يعزز من قوة ودقة التنبؤات النهائية. يتم تلخيص المنهجية بصريًا في الشكل 3، الذي يوضح التدفق من ترميز BERT إلى استخراج الميزات والتصنيف النهائي.
في النتائج التجريبية، تم تقسيم مجموعة البيانات إلى مجموعات تدريب (80%) واختبار (20%). تم تقييم أداء النموذج باستخدام مقاييس قياسية مثل الدقة، والدقة، والاسترجاع، التي تقيم التوافق بين الفئات الفعلية والمتوقعة. يتم تفصيل الدقة، التي تعرف على أنها نسبة الحالات المصنفة بشكل صحيح إلى العدد الإجمالي للحالات، في المعادلة 2. هذه المقاييس حاسمة لتقييم فعالية نماذج التعلم الآلي في هذا السياق.
نقاش
تسلط قسم النقاش في ورقة البحث الضوء على التقدم في تقنيات التعلم العميق للكشف عن مرض الزهايمر (AD) وتشخيصه. تم استخدام نماذج مختلفة، بما في ذلك الشبكات العصبية المتكررة (RNN)، وشبكات الذاكرة طويلة وقصيرة الأمد (LSTM)، والشبكات العصبية التلافيفية (CNN)، لتحليل اللغة المنطوقة والصور الطبية، مثل فحوصات التصوير بالرنين المغناطيسي، محققة معدلات دقة عالية. ومن الجدير بالذكر أن نموذج CNN أظهر دقة بنسبة 92% في التمييز بين فحوصات الدماغ الصحية والتنكسية. بالإضافة إلى ذلك، ظهرت طرق قائمة على المحولات كأدوات فعالة لتشخيص مرض الزهايمر من النص، مع نماذج مثل GP-Net التي حققت دقة قدرها 74.3% على مجموعة بيانات ADReSS و91.4% على مجموعة بيانات Pitt. تؤكد هذه النتائج على إمكانيات التعلم العميق لتسهيل التشخيصات المبكرة والأكثر موثوقية، فضلاً عن تحديد عوامل الخطر الجديدة لمرض الزهايمر.
علاوة على ذلك، يناقش البحث دمج الأساليب متعددة الوسائط، التي تجمع بين النص والصوت وبيانات التصوير العصبي لتعزيز دقة التشخيص. أظهرت تقنيات مثل CascadeNet لتحليل إشارات EEG والنماذج الهجينة التي تتضمن أنواع بيانات متنوعة نتائج واعدة، حيث حقق بعضها دقة تزيد عن 90% في تصنيف الخرف الآلي. ومع ذلك، لا تزال التحديات قائمة، بما في ذلك الحاجة إلى مجموعات بيانات أكبر لتحسين قابلية تعميم النموذج وقابلية تفسير نماذج التعلم العميق، التي غالبًا ما تعمل كـ “صناديق سوداء”. يؤكد المؤلفون على أهمية استمرار البحث لتحسين هذه المنهجيات واستكشاف تطبيقاتها السريرية، بهدف تحسين رعاية المرضى والنتائج في الكشف عن مرض الزهايمر وإدارته.
القيود
يسلط قسم القيود الضوء على عدة قيود حاسمة للنموذج الهجين الجماعي، الذي حقق دقة ملحوظة تبلغ 94.98% على مجموعة Cookie Theft باللغة الإنجليزية من مجموعة DementiaBank Pitt. أولاً، نطاق مجموعة البيانات محدود إلى مجموعة نصية واحدة باللغة الإنجليزية، مما يثير القلق بشأن قابلية تطبيق النموذج على عينات الكلام غير الإنجليزية بسبب الاختلافات المحتملة في الميزات اللغوية والثقافية. بالإضافة إلى ذلك، فإن قابلية تعميم التركيبة السكانية محدودة، حيث تتكون مجموعة Pitt بشكل أساسي من كبار السن الناطقين باللغة الإنجليزية من أمريكا الشمالية، مما قد لا يعكس اللهجات المتنوعة وطرق السرد الموجودة في مجموعات سكانية أخرى.
لمعالجة هذه القيود، يقترح المؤلفون اتجاهات بحث مستقبلية، بما في ذلك التقييمات عبر اللغات مع مجموعات بيانات مثل مجموعة Dem@Care الفرنسية ومجموعة Santiago الإسبانية. ستقوم هذه التقييمات بتقييم قوة النموذج وتحديد التعديلات اللازمة الخاصة باللغة. علاوة على ذلك، قد يكشف دمج التصنيف الديموغرافي في تقييمات النموذج عن التحيزات عبر مجموعات ثقافية مختلفة، مما يعزز العدالة التشخيصية ويكشف عن علامات سوسيولغوية تنبؤية لمرض الزهايمر. يقترح المؤلفون أيضًا استكشاف تحسينات محددة للمجال على BERT على مجموعات البيانات السريرية والمتعلقة بالخرف، بالإضافة إلى التحقيق في آليات قائمة على الانتباه لتحسين قابلية التفسير والتقاط التحولات الدلالية الدقيقة التي تشير إلى التدهور الإدراكي.
DOI: https://doi.org/10.1038/s41598-025-08362-y
PMID: https://pubmed.ncbi.nlm.nih.gov/40634379
Publication Date: 2025-07-09
Author(s): Sana Latif et al.
Primary Topic: Brain Tumor Detection and Classification
Overview
The research addresses the urgent need for early detection of Alzheimer’s disease (AD), a progressive neurodegenerative disorder that significantly impairs cognitive functions, particularly language comprehension. Current diagnostic methods are limited, and the study proposes a novel approach utilizing a two-branch deep learning model that combines a BERT encoder and a recurrent convolutional neural network (RCNN). The BERT encoder processes textual data from clinical notes, converting it into fixed-length vectors, which are then analyzed through convolutional and LSTM layers to predict AD presence. The RCNN also processes the same text data, and the outputs from both branches are fused using ensemble learning to enhance classification accuracy.
The proposed model was evaluated on the Cookie Theft subset of the DementiaBank Pitt Corpus, achieving an impressive accuracy of 94.98%, an F1-score of 0.9523, and an AUC of 0.93. These results indicate that the model outperforms existing state-of-the-art techniques for early AD diagnosis based on text analysis. This research highlights the potential of artificial intelligence, particularly deep learning algorithms, in improving diagnostic accuracy and facilitating timely intervention for Alzheimer’s disease.
Introduction
In this section, the authors present three significant innovations in their approach to detecting Alzheimer’s disease from clinical transcripts, building upon existing research that has utilized decision-level ensembles of BERT and CNN for text classification.
First, they introduce a hybrid architecture that integrates BERT, CNN, and LSTM branches, enabling the model to effectively capture both long-range dependencies and localized sequential patterns in language that may indicate cognitive decline. Second, they implement feature-level stacking with adaptive weighting, where 128-dimensional embeddings from the penultimate layer of each branch are concatenated and processed through a shallow multi-layer perceptron (MLP). This MLP learns dynamic weights for each feature, allowing the model to prioritize the most informative branches based on the specific input. Lastly, the authors propose a clinically-aware preprocessing pipeline that includes normalization of clinical concepts, word sense disambiguation, and correction of speech disfluencies, ensuring that relevant linguistic markers for Alzheimer’s are accurately preserved and highlighted.
Methods
The proposed methodology in this research involves a two-branch neural network model designed for the classification of Alzheimer’s Disease (AD) based on text data. The first branch utilizes a BERT encoder to transform the input text into a fixed-length vector, which is subsequently processed through a convolutional layer, an LSTM layer, and a fully connected layer to predict AD presence. The second branch employs a recurrent convolutional neural network (RCNN) that also processes the text data to produce classification outputs. These two branches are integrated using an ensemble learning approach, enhancing the robustness and accuracy of the final predictions. The methodology is visually summarized in Figure 3, illustrating the flow from BERT encoding to feature extraction and final classification.
In the experimental results, the dataset was divided into training (80%) and testing (20%) subsets. The model’s performance was evaluated using standard metrics such as accuracy, precision, and recall, which assess the correspondence between actual and predicted classes. The accuracy, defined as the ratio of correctly classified instances to the total number of instances, is detailed in Equation 2. These metrics are critical for evaluating the effectiveness of machine learning models in this context.
Discussion
The discussion section of the research paper highlights the advancements in deep learning techniques for the detection and diagnosis of Alzheimer’s disease (AD). Various models, including Recurrent Neural Networks (RNN), Long Short-Term Memory (LSTM) networks, and Convolutional Neural Networks (CNN), have been employed to analyze spoken language and medical images, such as MRI scans, achieving high accuracy rates. Notably, a CNN model demonstrated a 92% accuracy in distinguishing between healthy and degenerative brain scans. Additionally, transformer-based methods have emerged as effective tools for diagnosing AD from text, with models like GP-Net achieving accuracies of 74.3% on the ADReSS dataset and 91.4% on the Pitt dataset. These findings underscore the potential of deep learning to facilitate earlier and more reliable diagnoses, as well as to identify novel risk factors for AD.
Furthermore, the paper discusses the integration of multimodal approaches, combining text, audio, and neuroimaging data to enhance diagnostic accuracy. Techniques such as CascadeNet for EEG signal analysis and hybrid models incorporating various data types have shown promising results, with some achieving over 90% accuracy in automated dementia classification. However, challenges remain, including the need for larger datasets to improve model generalizability and the interpretability of deep learning models, which often function as “black boxes.” The authors emphasize the importance of continued research to refine these methodologies and explore their clinical applications, ultimately aiming to improve patient care and outcomes in Alzheimer’s disease detection and management.
Limitations
The section on limitations highlights several critical constraints of the hybrid ensemble model, which achieved a notable accuracy of 94.98% on the English Cookie Theft subset of the DementiaBank Pitt Corpus. Firstly, the dataset’s scope is limited to a single English-language corpus, raising concerns about the model’s applicability to non-English speech samples due to potential variations in linguistic and cultural features. Additionally, the demographic generalizability is restricted, as the Pitt Corpus primarily consists of older, English-speaking North Americans, which may not reflect the diverse dialects and storytelling conventions present in other populations.
To address these limitations, the authors suggest future research directions, including cross-lingual evaluations with datasets such as the French Dem@Care corpus and the Spanish-based Santiago Corpus. Such evaluations would assess the model’s robustness and identify necessary language-specific adaptations. Furthermore, incorporating demographic stratification in model evaluations could uncover biases across different cultural groups, enhancing diagnostic equity and revealing sociolinguistic markers predictive of Alzheimer’s Disease. The authors also propose exploring domain-specific fine-tuning of BERT on clinical and dementia-related corpora, as well as investigating attention-based mechanisms to improve interpretability and capture subtle semantic shifts indicative of cognitive decline.
