DOI: https://doi.org/10.3389/fmed.2025.1635796
PMID: https://pubmed.ncbi.nlm.nih.gov/40917831
تاريخ النشر: 2025-08-22
المؤلف: Abdullah Almuhaimeed وآخرون
الموضوع الرئيسي: كشف وتصنيف أورام الدماغ
نظرة عامة
تقدم ورقة البحث نموذجًا جديدًا للتعلم العميق يدمج بين Swin Transformer وشبكة الخصومة المشروطة المعززة (AE-cGAN) لمعالجة تعقيدات تصنيف أورام الدماغ. غالبًا ما تتجاهل الطرق التقليدية الميزات الحرجة بسبب عدم توازن البيانات وقدرات استخراج الميزات المحدودة. يعزز النموذج المقترح تنوع مجموعة البيانات من خلال توليد الصور الاصطناعية عبر AE-cGAN، بينما يلتقط Swin Transformer بشكل فعال الاعتماديات المحلية والعالمية في البيانات. حقق النموذج معدلات دقة ملحوظة بلغت 99.54% و98.9% على مجموعتي بيانات Figshare وKaggle، على التوالي، متجاوزًا التقنيات الحديثة الموجودة.
تؤكد النتائج على قدرة النموذج في التخفيف من عدم توازن البيانات وتحسين استخراج الميزات، مما يؤدي إلى أداء تصنيفي متفوق. ستركز الأبحاث المستقبلية على تكييف النموذج للتطبيقات السريرية، مما يمكّن التقييمات في الوقت الحقيقي، ودمج بيانات التصوير متعددة الأنماط مثل مسحات MRI T1 وT2 وFLAIR. بالإضافة إلى ذلك، ستوجه الجهود نحو تحسين جودة البيانات الاصطناعية التي يتم توليدها بواسطة AE-cGAN، واستكشاف التعلم الانتقالي، وضمان قابلية تفسير النموذج وقوته في البيئات السريرية. الهدف النهائي هو تحسين عملية التشخيص لأورام الدماغ، مما يحسن كل من السرعة والدقة في البيئات السريرية.
مقدمة
تسلط مقدمة ورقة البحث الضوء على التقدم الكبير في التصوير الطبي والذكاء الحسابي، لا سيما من خلال تطبيق التعلم العميق، والاندماج متعدد الأنماط، والأساليب المستوحاة من الكم. لقد حسنت هذه الابتكارات بشكل ملحوظ النتائج السريرية في مجالات طبية متنوعة، مثل تحسين الأساليب الجراحية للنزيف داخل الدماغ وتطوير أطر عمل تعتمد على EEG لتقدير الحالة العقلية والتعرف على المشاعر. لقد عالج دمج مشفرات تلقائية مقنعة متعددة الأنماط ووحدات خفيفة الوزن التحديات في استخراج الميزات وكفاءة الحساب، لا سيما في تصنيف الأمراض والتصوير الفائق الدقة. علاوة على ذلك، أظهرت الاستراتيجيات المستوحاة بيولوجيًا والشبكات المدفوعة بالمحولين وعدًا في تحسين دقة التشخيص عبر تطبيقات متنوعة، بما في ذلك التقييمات النفسية العصبية وتصوير الأسنان.
تؤكد الورقة على الحاجة الملحة لأنظمة تشخيص متقدمة في مجال تصنيف أورام الدماغ المعقد والعالي المخاطر، حيث تكافح الطرق التشخيصية التقليدية غالبًا بسبب العرض غير المتجانس للأورام والأعراض المبكرة الدقيقة. ظهرت أنظمة الكشف والتشخيص المدعومة بالكمبيوتر (CADe وCADx)، التي تستفيد من الذكاء الاصطناعي ورؤية الكمبيوتر، كأدوات أساسية لأطباء الأشعة، مما يعزز دقة وكفاءة التشخيص. على الرغم من التقدم المحرز، لا تزال التحديات قائمة، لا سيما فيما يتعلق بتوافر مجموعات بيانات موثقة عالية الجودة لتدريب نماذج التعلم العميق. أظهرت إدخال تقنيات تعلم الآلة المختلفة، بما في ذلك الشبكات العصبية التلافيفية (CNNs) والنماذج الهجينة، وعدًا في تحسين دقة التصنيف، ومع ذلك لا تزال القضايا المتعلقة بالعبء الحسابي والحاجة إلى أداء في الوقت الحقيقي في البيئات السريرية قائمة. تظهر التطورات الأخيرة، مثل GAN التعلم الكامن المتباين ذو التدفق المزدوج ومحول الرؤية غير القابل للدوران، الإمكانية لتعزيز القدرات التشخيصية بشكل أكبر مع معالجة قيود الطرق التقليدية.
الطرق
في هذه الدراسة، تم استخدام مجموعتين من بيانات MRI المتاحة للجمهور للتحقق من صحة نموذج شبكة الخصومة التوليدية (GAN) المقترح لتصنيف أورام الدماغ. تتكون مجموعة البيانات الأولى، مجموعة بيانات MRI للدماغ من Kaggle، من أكثر من 3,000 صورة MRI مصنفة إلى أربع فئات: الورم الدبقي، ورم السحايا، ورم الغدة النخامية، وعدم وجود ورم. تم تنظيم مجموعة البيانات باستخدام أخذ عينات طبقية للحفاظ على نسب الفئات وتم تقسيمها إلى مجموعات تدريب (80%)، والتحقق (10%)، والاختبار (10%). تحتوي مجموعة البيانات الثانية، مجموعة بيانات أورام الدماغ من Figshare، على 3,064 مسح MRI بالأبيض والأسود عبر ثلاثة أنواع من الأورام: الورم الدبقي، ورم الغدة النخامية، ورم السحايا، وتم تنظيمها أيضًا باستخدام أخذ عينات طبقية بنفس النسب.
تم إجراء التجارب على Google Colab Pro، مستفيدين من وحدة معالجة الرسوميات NVIDIA Tesla T4 وبيئة افتراضية لتحقيق الكفاءة الحسابية. تم ضبط المعلمات الفائقة تجريبيًا، حيث استخدم المشفر التلقائي مُحسِّن Adam ودالة خسارة متوسط مربع الخطأ (MSE)، بينما استخدم GAN المشروط (cGAN) مُحسِّن Adam بمعدل تعلم قدره \(1 \times 10^{-4}\) وخسارة cGAN القياسية. استخدم مصنف Swin Transformer مُحسِّن AdamW بمعدل تعلم قدره \(3 \times 10^{-5}\) وخسارة الانتروبيا المتقاطعة. لمعالجة عدم توازن الفئات، تم توليد صور اصطناعية باستخدام cGAN، مما أسفر عن 2,000 صورة إضافية لمجموعة البيانات الأولى و1,500 لمجموعة البيانات الثانية. اختلفت أوقات التدريب، حيث استغرق المشفر التلقائي حوالي 215 ثانية لـ 15 دورة، وGAN حوالي 8,200 ثانية لـ 100 دورة، والمصنف حوالي 5,493 ثانية عبر ثلاث طيات من التحقق المتقاطع.
المناقشة
تقدم الدراسة نموذجًا هجينًا جديدًا للتعلم العميق (DL)، يُطلق عليه نموذج GAS، والذي يدمج بين المشفرات التلقائية (AE) لاستخراج الميزات، وشبكات الخصومة التوليدية المشروطة (cGAN) لتوليد البيانات الاصطناعية، وSwin Transformers لتصنيف أورام الدماغ. تشمل الأهداف الرئيسية معالجة عدم توازن البيانات ومجموعات البيانات الموثقة المحدودة من خلال توليد بيانات اصطناعية عبر cGAN، مما يعزز من قوة النموذج وقدراته على التعميم. تم تقييم النموذج على مجموعات بيانات أورام الدماغ المتاحة للجمهور، مما أظهر دقة متفوقة مقارنة بالطرق الحديثة، حيث حقق دقة إجمالية بلغت 99.54% على مجموعة البيانات الأولى و98.9% على مجموعة البيانات الثانية.
تستخدم البنية مشفرًا تلقائيًا لتعلم مساحة كامنة مضغوطة من صور MRI الحقيقية، مما يساعد في التمييز بين الصور الحقيقية والاصطناعية أثناء التصنيف. يقوم cGAN بتوليد صور MRI اصطناعية واقعية مشروطة على تسميات نوع الورم، مما يزيد من مجموعة البيانات ويحسن أداء التصنيف. يُعرف Swin Transformer بآلية الانتباه الهرمية الخاصة به، حيث يلتقط بشكل فعال الميزات المحلية والعالمية من البيانات عالية الأبعاد، مما يسهل تصنيف الأورام بدقة. تشير النتائج إلى أن النموذج يتفوق في تمييز أنواع الأورام، لا سيما الورم الدبقي ورم السحايا، على الرغم من بعض التحديات بسبب التشابهات البصرية. بشكل عام، تؤكد النتائج على إمكانيات نموذج GAS للتطبيقات السريرية في التشخيص الطبي، لا سيما في تعزيز دقة وكفاءة تصنيف أورام الدماغ.
DOI: https://doi.org/10.3389/fmed.2025.1635796
PMID: https://pubmed.ncbi.nlm.nih.gov/40917831
Publication Date: 2025-08-22
Author(s): Abdullah Almuhaimeed et al.
Primary Topic: Brain Tumor Detection and Classification
Overview
The research paper presents a novel deep learning model that integrates a Swin Transformer with an Adversarially Enhanced Conditional Generative Adversarial Network (AE-cGAN) to address the complexities of brain tumor classification. Traditional methods often overlook critical features due to data imbalance and limited feature extraction capabilities. The proposed model enhances dataset diversity through synthetic image generation via AE-cGAN, while the Swin Transformer effectively captures both local and global dependencies in the data. The model achieved remarkable accuracy rates of 99.54% and 98.9% on the Figshare and Kaggle datasets, respectively, surpassing existing state-of-the-art techniques.
The findings underscore the model’s ability to mitigate data imbalance and improve feature extraction, leading to superior classification performance. Future research will focus on adapting the model for clinical applications, enabling real-time assessments, and integrating multimodal imaging data such as T1, T2, and FLAIR MRI scans. Additionally, efforts will be directed towards enhancing the quality of synthetic data generated by the AE-cGAN, exploring transfer learning, and ensuring the model’s interpretability and robustness in clinical settings. The ultimate aim is to refine the diagnostic process for brain tumors, improving both speed and accuracy in clinical environments.
Introduction
The introduction of the research paper highlights significant advancements in medical imaging and computational intelligence, particularly through the application of deep learning, multimodal fusion, and quantum-inspired methods. These innovations have notably improved clinical outcomes in various medical domains, such as enhancing surgical approaches for intracerebral hemorrhage and advancing EEG-based frameworks for mental state estimation and emotion recognition. The integration of multimodal masked autoencoders and lightweight modules has addressed challenges in feature extraction and computational efficiency, particularly in disease staging and super-resolution imaging. Furthermore, biologically inspired strategies and transformer-driven networks have shown promise in improving diagnostic accuracy across diverse applications, including neuropsychological assessments and dental imaging.
The paper emphasizes the critical need for advanced diagnostic systems in the complex and high-risk domain of brain tumor classification, where traditional diagnostic methods often struggle due to the heterogeneous presentation of tumors and subtle early-stage symptoms. Computer-assisted detection and diagnosis (CADe and CADx) systems, leveraging artificial intelligence and computer vision, have emerged as essential tools for radiologists, enhancing diagnostic accuracy and efficiency. Despite the progress made, challenges remain, particularly concerning the limited availability of high-quality annotated datasets for training deep learning models. The introduction of various machine learning techniques, including convolutional neural networks (CNNs) and hybrid models, has shown promise in improving classification accuracy, yet issues related to computational burden and the need for real-time performance in clinical settings persist. Recent advancements, such as the Dual-Stream Contrastive Latent Learning GAN and the Rotation Invariant Vision Transformer, demonstrate the potential for further enhancing diagnostic capabilities while addressing the limitations of traditional methods.
Methods
In this study, two publicly available MRI datasets were utilized to validate a proposed Generative Adversarial Network (GAN) model for brain tumor classification. Dataset I, the Kaggle Brain MRI Dataset, comprises over 3,000 labeled MRI images categorized into four classes: glioma, meningioma, pituitary tumor, and no tumor. The dataset was organized using stratified sampling to maintain class proportions and was split into training (80%), validation (10%), and testing (10%) sets. Dataset II, the Figshare Brain Tumor Data Set, contains 3,064 grayscale MRI scans across three tumor types: glioma, pituitary tumor, and meningioma, also organized with stratified sampling in the same proportions.
The experiments were conducted on Google Colab Pro, leveraging an NVIDIA Tesla T4 GPU and a virtualized environment for computational efficiency. Hyperparameters were empirically tuned, with the autoencoder using the Adam optimizer and Mean Squared Error (MSE) loss function, while the conditional GAN (cGAN) employed the Adam optimizer with a learning rate of \(1 \times 10^{-4}\) and standard cGAN loss. The Swin Transformer classifier utilized the AdamW optimizer with a learning rate of \(3 \times 10^{-5}\) and Cross-Entropy Loss. To address class imbalance, synthetic images were generated using the cGAN, resulting in 2,000 additional images for Dataset I and 1,500 for Dataset II. Training times varied, with the autoencoder taking approximately 215 seconds for 15 epochs, the GAN around 8,200 seconds for 100 epochs, and the classifier approximately 5,493 seconds across three folds of cross-validation.
Discussion
The study presents a novel hybrid deep learning (DL) model, termed the GAS model, which integrates Autoencoders (AE) for feature extraction, Conditional Generative Adversarial Networks (cGAN) for synthetic data generation, and Swin Transformers for the classification of brain tumors. The primary objectives include addressing data imbalance and limited annotated datasets by generating synthetic data through cGAN, thereby enhancing the model’s robustness and generalization capabilities. The model was evaluated on publicly available brain tumor datasets, demonstrating superior accuracy compared to state-of-the-art methods, achieving an overall accuracy of 99.54% on Dataset I and 98.9% on Dataset II.
The architecture employs an autoencoder to learn a compact latent space from real MRI images, which aids in distinguishing between real and synthetic images during classification. The cGAN generates realistic synthetic MRI images conditioned on tumor type labels, augmenting the dataset and improving classification performance. The Swin Transformer, known for its hierarchical attention mechanism, effectively captures local and global features from high-dimensional data, facilitating accurate tumor classification. The results indicate that the model excels in differentiating tumor types, particularly glioma and meningioma, despite some challenges due to visual similarities. Overall, the findings underscore the potential of the GAS model for clinical applications in medical diagnosis, particularly in enhancing the accuracy and efficiency of brain tumor classification.
