DOI: https://doi.org/10.35378/gujs.1684696
تاريخ النشر: 2026-02-15
المؤلف: Faruk Özger وآخرون
الموضوع الرئيسي: كشف وتصنيف أورام الدماغ
نظرة عامة
تقوم هذه الدراسة بتقييم منهجي لعشر هياكل بارزة في التعلم العميق، تقارن بين خمسة شبكات عصبية تلافيفية (CNNs) وخمسة محولات رؤية (ViTs) لتصنيف أورام الدماغ باستخدام بيانات التصوير بالرنين المغناطيسي. تشير النتائج إلى أن كلا عائلتي النماذج تحققان دقة تصنيف عالية، حيث تحقق Inception v4 أعلى دقة بنسبة 96.73%، مما يظهر فعالية المعالجة التلافيفية متعددة المقاييس. يتبعها محول PiT-Base بدقة 96.41%، مع استخدام عدد أقل من المعلمات، مما يبرز مزايا تصميمات تجميع الرموز الفعالة.
تشدد الدراسة على أن التعقيد المعماري أكثر أهمية من العدد الصريح للمعلمات في تحديد أداء النموذج، إلى جانب أهمية سلامة البيانات وبروتوكولات التدريب القوية. تشير الدقة التشخيصية العالية لهذه النماذج إلى إمكاناتها كنظم دعم القرار في البيئات السريرية، مما يسهل تصنيف المرضى واستراتيجيات العلاج. لتعزيز الاعتماد السريري، تدعو الدراسة إلى دمج طرق التفسير، والتحقق من صحة النماذج على مجموعات بيانات خارجية كبيرة، وتنفيذ قياس عدم اليقين لتحسين موثوقية النموذج وثقة الأطباء. معالجة هذه المجالات أمر ضروري لتطوير هذه الهياكل عالية الأداء إلى أدوات سريرية حيوية في علم الأورام العصبية.
مقدمة
تسلط مقدمة هذه الورقة البحثية الضوء على التحدي الكبير الذي تطرحه أورام الدماغ، والتي تنتج عن تكاثر غير طبيعي للخلايا وتؤدي إلى معدلات مرضية ووفيات كبيرة في جميع أنحاء العالم، مع تشخيص أكثر من 300,000 حالة جديدة سنويًا. تظهر الأورام العدوانية مثل الورم الدبقي معدلات بقاء منخفضة بشكل خاص لمدة خمس سنوات، مما يبرز الحاجة الملحة إلى تشخيص دقيق وفي الوقت المناسب لتحسين نتائج المرضى. بينما يُعتبر التصوير بالرنين المغناطيسي (MRI) المعيار الذهبي لتصوير الأورام، فإن التحليل اليدوي لهذه المسحات يتطلب جهدًا كبيرًا وعرضة للاختلافات، مما يستدعي تطوير أدوات تشخيص آلية.
لمعالجة هذه الحاجة، تستكشف الورقة تطبيق الذكاء الاصطناعي (AI)، وخاصة التعلم العميق، في تحليل الصور الطبية. تؤكد على قدرات الشبكات العصبية التلافيفية (CNNs) ومحولات الرؤية (ViTs) في اكتشاف الأنماط المعقدة داخل مسحات التصوير بالرنين المغناطيسي. على الرغم من النتائج الواعدة للنماذج الفردية، إلا أن دراسة مقارنة شاملة تقيم أحدث هياكل CNN وViT على مهمة تصنيف أورام الدماغ الموحدة كانت مفقودة. تهدف هذه الدراسة إلى سد هذه الفجوة من خلال إجراء تحليل مقارن لعشر نماذج متقدمة—خمسة CNNs وخمسة ViTs—باستخدام مجموعة بيانات عامة من صور الرنين المغناطيسي للدماغ. تقيم الدراسة أداء النموذج من خلال مقاييس مختلفة، وتفحص التوازن بين الدقة والكفاءة الحسابية، وتستخدم تقنيات الذكاء الاصطناعي القابل للتفسير (XAI) لتعزيز شفافية عمليات اتخاذ القرار للنماذج.
الطرق
تحدد قسم المنهجية في هذه الدراسة تطبيق الذكاء الاصطناعي في تشخيص أورام الدماغ، موضحة مجموعات البيانات، وتقنيات التعلم العميق، والأساليب المحددة المستخدمة، بما في ذلك الشبكات العصبية التلافيفية (CNNs)، التعلم الانتقالي، محولات الرؤية (ViT)، وطرق تعزيز البيانات.
تم إجراء التصميم التجريبي على محطة عمل عالية الأداء تعمل بنظام Ubuntu 24.04 LTS، تحتوي على معالج Intel Core i9 من الجيل الرابع عشر وبطاقة رسومات NVIDIA GeForce RTX 4090 مع 24 جيجابايت من ذاكرة GDDR6X VRAM، بالإضافة إلى 64 جيجابايت من ذاكرة DDR5 RAM. سهلت هذه الإعدادات تنفيذ المهام الحسابية المكثفة، حيث تم بناء بيئة البرمجيات على Python 3.12 وPyTorch 2.5، باستخدام NVIDIA CUDA 12.4 وcuDNN 9.1 لتسريع GPU. لضمان إمكانية التكرار، تم إجراء جميع تدريبات وتقييمات النموذج تحت ظروف متسقة ضمن هذا الإطار الحسابي.
النتائج
يقدم قسم النتائج تحليلًا مقارنًا لخمس هياكل من الشبكات العصبية التلافيفية (CNN)—Xception، Inceptionv4، ConvNeXt-Base، InceptionNeXt-Base، وEfficientNetV2-Medium—المطبقة على تصنيف أورام الدماغ. برز Inceptionv4 كأفضل أداء، حيث حقق دقة 0.9673، ودقة 0.9660، واسترجاع 0.9704، ودرجة F1 0.9681، مما يظهر فعالية التصميمات متعددة المقاييس. تبعها ConvNeXt-Base بدقة 0.9641 ولكن كان لديها عدد معلمات أكبر بكثير يبلغ 87.57 مليون، مما يشير إلى أن الكفاءة المعمارية أمر حاسم. قدم EfficientNetV2-Medium أداءً متوازنًا بدقة 0.9542، بينما كان Xception، على الرغم من كونه النموذج الأخف بواقع 20.81 مليون معلمة، لديه أدنى دقة بنسبة 0.9461، مما يبرز التوازن بين التكلفة الحسابية والقوة التنبؤية.
في فئة محول الرؤية (ViT)، تم تقييم خمسة نماذج—ViT-Base-Patch16، Swin-Base، BeiT-Base، DeiT-Base، وPiT-Base. كان PiT-Base الأكثر فعالية، حيث حقق أعلى دقة بنسبة 0.9641 ودرجة F1 0.9638 بينما كان الأخف بواقع 72.74 مليون معلمة. كان Swin-Base، النموذج الأثقل بواقع 86.75 مليون معلمة، أيضًا أداءً جيدًا بدقة 0.9624. تفوق DeiT-Base في الدقة، مسجلاً 0.9644، بينما قدم BeiT-Base وViT-Base-Patch16 أداءً قويًا ولكن أقل نسبيًا. تؤكد النتائج على أهمية كل من بنية النموذج وكفاءة المعلمات في تحقيق دقة تنبؤية عالية في مهام تصنيف أورام الدماغ.
المناقشة
تسلط المناقشة الضوء على التقدم الكبير في اكتشاف وتصنيف أورام الدماغ من خلال التعلم العميق، بينما تتناول أيضًا التحديات المستمرة. على الرغم من الدقة المثيرة للإعجاب التي حققتها نماذج مختلفة، فإن مشكلات مثل مجموعات البيانات المحدودة وغير المتوازنة تعيق القوة، مما يستدعي جمع بيانات أكبر وأكثر تنوعًا. تعتمد العديد من الدراسات على مجموعات بيانات التصوير بالرنين المغناطيسي المتاحة للجمهور التي قد لا تمثل بشكل كافٍ تعقيدات مجموعات المرضى المتنوعة. بالإضافة إلى ذلك، تثير الطبيعة الغامضة لعمليات اتخاذ القرار في التعلم العميق مخاوف بشأن موثوقية النموذج في التطبيقات الطبية، مما يبرز الحاجة إلى تحسين القابلية للتفسير. يُقترح دمج المدخلات متعددة الوسائط، بما في ذلك البيانات الجينية والأنسجة، كوسيلة لتعزيز الدقة التشخيصية.
أظهرت الأبحاث الحديثة فعالية النماذج الهجينة في التعلم العميق، التي تجمع بين عدة شبكات مدربة مسبقًا وتقنيات تحسين لتحقيق دقة تصنيف عالية لأورام الدماغ. على سبيل المثال، حقق نموذج هجيني يستخدم GoogleNet وAlexNet وغيرهما معدل دقة ملحوظ من خلال استراتيجية التصويت بالأغلبية. علاوة على ذلك، أظهرت التقدمات في الهياكل مثل محولات الرؤية (ViTs) وEfficientNet وعدًا في تحسين أداء التصنيف مع الحفاظ على الكفاءة الحسابية. تستفيد ViTs، بشكل خاص، من آليات الانتباه الذاتي لالتقاط الاعتماديات بعيدة المدى، مما يعزز فعاليتها في مهام التصوير الطبي. بشكل عام، يمثل دمج التعلم العميق مع تقنيات تحسين متقدمة ومصادر بيانات متنوعة مسارًا واعدًا لتحسين دقة وموثوقية تشخيص أورام الدماغ.
القيود
تقدم الدراسة مقارنة شاملة لنماذج التعلم العميق لتحليل صور الرنين المغناطيسي المعززة بالتباين T1؛ ومع ذلك، فإنها مقيدة بعدة قيود تبرز الاتجاهات المحتملة للبحث المستقبلي. أولاً، يقتصر الاستخدام الحصري لتسلسل تصوير بالرنين المغناطيسي واحد على دمج معلومات قيمة من أنماط أخرى، مثل T2 وFLAIR. يجب أن تهدف التحقيقات المستقبلية إلى دمج بيانات الرنين المغناطيسي متعددة الأنماط، حيث أظهر ذلك تحسينًا كبيرًا في الدقة التشخيصية. ثانيًا، فإن الاعتماد على مجموعة بيانات عامة واحدة يعرض مخاطر التحيز ويحد من قابلية تعميم النماذج على الصور من مراكز طبية متنوعة. يجب أن تعطي الأعمال المستقبلية الأولوية للتحقق الخارجي عبر مجموعات متعددة المراكز والتركيز على تجميع مجموعات بيانات أكبر وأكثر تنوعًا لتحسين قوة النموذج ضد التغيرات الديموغرافية والتقنية.
بالإضافة إلى ذلك، تقوم المنهجية الحالية بتحليل كل شريحة محورية بشكل مستقل، مما، على الرغم من كفاءته الحسابية، يتجاهل الخصائص الحجمية ثلاثية الأبعاد لصور الرنين المغناطيسي. يجب أن تستكشف الأبحاث المستقبلية الهياكل ثلاثية الأبعاد الحقيقية التي يمكن أن تعالج الأحجام الكاملة، مما يستفيد من الاستمرارية بين الشرائح للحصول على رؤى سياقية أغنى. أخيرًا، يركز التركيز التشخيصي للدراسة على استبعاد فئة “لا ورم”، مما يحد من فائدة النموذج كأداة فحص مستقلة. منطقة حاسمة للتطوير المستقبلي هي إنشاء نظام من مرحلتين يكشف أولاً عن الشذوذات (ورم مقابل لا ورم) قبل تصنيف الآفات المحددة، مما يعالج كل من الدقة والسلامة في التطبيقات السريرية.
DOI: https://doi.org/10.35378/gujs.1684696
Publication Date: 2026-02-15
Author(s): Faruk Özger et al.
Primary Topic: Brain Tumor Detection and Classification
Overview
This research systematically evaluates ten prominent deep learning architectures, comparing five convolutional neural networks (CNNs) and five Vision Transformers (ViTs) for brain tumor classification using MRI data. The findings indicate that both model families achieve high classification accuracy, with Inception v4 attaining the highest accuracy of 96.73%, demonstrating the effectiveness of multiscale convolutional processing. The PiT-Base transformer closely follows with an accuracy of 96.41%, while utilizing fewer parameters, highlighting the advantages of efficient token-pooling designs.
The study emphasizes that architectural sophistication is more critical than the sheer number of parameters in determining model performance, alongside the importance of data integrity and robust training protocols. The high diagnostic accuracy of these models suggests their potential as decision-support systems in clinical settings, facilitating patient triage and treatment strategies. To enhance clinical adoption, the research advocates for the integration of interpretability methods, validation on large-scale external datasets, and the implementation of uncertainty quantification to improve model reliability and clinician trust. Addressing these areas is essential for evolving these high-performing architectures into vital clinical tools in neuro-oncology.
Introduction
The introduction of this research paper highlights the significant challenge posed by brain tumors, which result from abnormal cell proliferation and lead to considerable morbidity and mortality worldwide, with over 300,000 new cases diagnosed annually. Aggressive tumors like glioblastoma exhibit particularly low five-year survival rates, underscoring the critical need for accurate and timely diagnosis to improve patient outcomes. While Magnetic Resonance Imaging (MRI) is the gold standard for tumor visualization, the manual analysis of these scans is labor-intensive and prone to inconsistencies, necessitating the development of automated diagnostic tools.
To address this need, the paper explores the application of artificial intelligence (AI), particularly deep learning, in medical image analysis. It emphasizes the capabilities of Convolutional Neural Networks (CNNs) and Vision Transformers (ViTs) in detecting complex patterns within MRI scans. Despite the promising results of individual models, a comprehensive comparative study assessing the latest CNN and ViT architectures on a standardized brain tumor classification task has been lacking. This research aims to fill that gap by conducting a comparative analysis of ten advanced models—five CNNs and five ViTs—using a public dataset of brain MRIs. The study evaluates model performance through various metrics, examines the trade-off between accuracy and computational efficiency, and employs Explainable AI (XAI) techniques to enhance the transparency of the models’ decision-making processes.
Methods
The methodology section of this study outlines the application of artificial intelligence in diagnosing brain tumors, detailing the datasets, deep learning techniques, and specific approaches employed, including convolutional neural networks (CNNs), transfer learning, Vision Transformers (ViT), and data augmentation methods.
The experimental design was conducted on a high-performance workstation running Ubuntu 24.04 LTS, featuring a 14th-generation Intel Core i9 processor and an NVIDIA GeForce RTX 4090 GPU with 24GB of GDDR6X VRAM, alongside 64GB of DDR5 RAM. This setup facilitated the execution of computationally intensive tasks, with the software environment built on Python 3.12 and PyTorch 2.5, utilizing NVIDIA CUDA 12.4 and cuDNN 9.1 for GPU acceleration. To ensure reproducibility, all model training and evaluation were performed under consistent conditions within this computational framework.
Results
The results section presents a comparative analysis of five Convolutional Neural Network (CNN) architectures—Xception, Inceptionv4, ConvNeXt-Base, InceptionNeXt-Base, and EfficientNetV2-Medium—applied to brain tumor classification. Inceptionv4 emerged as the top performer, achieving an accuracy of 0.9673, precision of 0.9660, recall of 0.9704, and an F1-score of 0.9681, demonstrating the effectiveness of multi-scale designs. ConvNeXt-Base followed closely with an accuracy of 0.9641 but had a significantly larger parameter count of 87.57 million, indicating that architectural efficiency is crucial. EfficientNetV2-Medium provided a balanced performance with an accuracy of 0.9542, while Xception, despite being the lightest model at 20.81 million parameters, had the lowest accuracy of 0.9461, highlighting the trade-off between computational cost and predictive power.
In the Vision Transformer (ViT) category, five models—ViT-Base-Patch16, Swin-Base, BeiT-Base, DeiT-Base, and PiT-Base—were evaluated. PiT-Base was the most effective, achieving the highest accuracy of 0.9641 and F1-score of 0.9638 while being the lightest at 72.74 million parameters. Swin-Base, the heaviest model at 86.75 million parameters, also performed well with an accuracy of 0.9624. DeiT-Base excelled in precision, recording 0.9644, while BeiT-Base and ViT-Base-Patch16 provided solid but comparatively lower performances. The findings underscore the importance of both model architecture and parameter efficiency in achieving high predictive accuracy in brain tumor classification tasks.
Discussion
The discussion highlights significant advancements in brain tumor detection and classification through deep learning, while also addressing ongoing challenges. Despite the impressive accuracy achieved by various models, issues such as limited and imbalanced datasets hinder robustness, necessitating larger and more diverse data collections. Many studies rely on publicly available MRI datasets that may not adequately represent the complexities of diverse patient populations. Additionally, the opaque nature of deep learning decision-making raises concerns regarding model reliability in medical applications, underscoring the need for improved explainability. The integration of multi-modal inputs, including genetic and histopathological data, is suggested as a means to enhance diagnostic accuracy.
Recent research has demonstrated the efficacy of hybrid deep learning models, which combine multiple pre-trained networks and optimization techniques to achieve high classification accuracy for brain tumors. For instance, a hybrid model utilizing GoogleNet, AlexNet, and others achieved a remarkable accuracy rate through a majority voting strategy. Furthermore, advancements in architectures such as Vision Transformers (ViTs) and EfficientNet have shown promise in improving classification performance while maintaining computational efficiency. ViTs, in particular, leverage self-attention mechanisms to capture long-range dependencies, enhancing their effectiveness in medical imaging tasks. Overall, the integration of deep learning with advanced optimization techniques and diverse data sources presents a promising pathway for improving the precision and reliability of brain tumor diagnosis.
Limitations
The study presents a thorough comparison of deep learning models for analyzing contrast-enhanced T1-weighted MRI images; however, it is constrained by several limitations that highlight potential directions for future research. Firstly, the exclusive use of a single MRI sequence limits the incorporation of valuable information from other modalities, such as T2 and FLAIR. Future investigations should aim to integrate multi-modal MRI data, as this has been shown to enhance diagnostic accuracy significantly. Secondly, reliance on a single public dataset poses risks of bias and limits the generalizability of the models to images from diverse medical centers. Future work should prioritize external validation across multicenter cohorts and focus on aggregating larger, more varied datasets to improve model robustness against demographic and technical variations.
Additionally, the current methodology analyzes each axial slice independently, which, while computationally efficient, neglects the 3D volumetric characteristics of MRI scans. Future research should explore true 3D architectures that can process entire volumes, thereby leveraging inter-slice continuity for richer contextual insights. Lastly, the study’s diagnostic focus excludes a ‘no-tumor’ class, limiting the model’s utility as a standalone screening tool. A critical area for future development is the creation of a two-stage system that first detects abnormalities (tumor vs. no-tumor) before classifying identified lesions, addressing both accuracy and safety in clinical applications.
