DOI: https://doi.org/10.2463/mrms.mp.2025-0107
PMID: https://pubmed.ncbi.nlm.nih.gov/41672496
تاريخ النشر: 2026-01-01
المؤلف: Binghua Li وآخرون
الموضوع الرئيسي: كشف وتصنيف أورام الدماغ
نظرة عامة
تقدم البحث شبكة موثوقة وقابلة للتفسير (MoRE-Net)، وهي إطار مبتكر مصمم لتعزيز قوة ووضوح نماذج التشخيص الطبي متعددة الوسائط، وخاصة لتصنيف أورام الدماغ في ظل ظروف فقدان الوسائط. يستخدم النموذج مشفرات مستقلة لكل وسيلة لاستخراج ميزات محددة لكل وسيلة ويشمل استراتيجيتين رئيسيتين: مشفرات قائمة على Mamba لنمذجة السياق العالمي بكفاءة ومعلم متعدد الوسائط عبر الإنترنت يستخدم خسارة المحاذاة خلال التدريب المبكر، والتي يتم التخلص منها تدريجياً.
تم تقييم MoRE-Net على مجموعة بيانات BraTS2020 مع 369 موضوعًا، حيث حقق دقة متوازنة متوسطة (BAC) تبلغ 73.5% ودقة تنشيط (AP) تبلغ 61.2%، متفوقًا على الطرق الأساسية بحوالي 15% و21%، على التوالي. أكدت التحقق على مجموعة بيانات ReMIND والدراسات الاستبعادية فعالية الاستراتيجيات المقترحة والموثوقية العامة للنموذج. تشير النتائج إلى أن MoRE-Net لا يتفوق فقط في دقة التشخيص ولكن أيضًا يحافظ على القابلية للتفسير، مما يبرز إمكانيته للتطبيق السريري في البيئات الطبية ذات المخاطر العالية. ستهدف الأبحاث المستقبلية إلى توسيع تطبيقاته ودمج تقدير عدم اليقين لتعزيز موثوقية السريرية.
مقدمة
تسلط مقدمة الورقة الضوء على الدور المتزايد للذكاء الاصطناعي (AI) في تحليل الصور الطبية، خاصة في التطبيقات التشخيصية والتنبؤية. بينما أظهر الذكاء الاصطناعي إمكانات كبيرة، فإن ضمان موثوقية هذه الأنظمة أمر حاسم، خاصة في البيئات السريرية. متطلبات أساسية للموثوقية هي القابلية للتفسير – حيث يتم توضيح تفكير النموذج من خلال أدوات مثل خرائط النسبة والتحليلات المضادة – والموثوقية، التي تشير إلى قدرة النموذج على الأداء بشكل موثوق تحت ظروف مختلفة، بما في ذلك الاضطرابات المدخلة وتحولات التوزيع.
تركز الورقة على تحليل أورام الدماغ ثلاثية الأبعاد المستندة إلى التصوير بالرنين المغناطيسي، والتي تتضمن مهامًا معقدة تتطلب كل من القابلية للتفسير والموثوقية. على الرغم من أن الدراسات السابقة قد حققت تقدمًا في تعزيز إما القابلية للتفسير أو الموثوقية، إلا أن هناك فجوة ملحوظة في الأبحاث التي تعالج كلا الجانبين في وقت واحد، خاصة في سياق تصنيف أورام الدماغ ثلاثية الأبعاد مع الوسائط المفقودة. يقترح المؤلفون تحسين دقة التشخيص والقابلية للتفسير مع ضمان عدم تأثير تقنيات الموثوقية على قدرة النموذج على التقاط الاعتماديات بعيدة المدى عبر الحجم بالكامل، وهو أمر حاسم لمهام التصنيف العالمية.
طرق
في هذه الدراسة، يقدم المؤلفون شبكة موثوقة وقابلة للتفسير (MoRE-Net)، وهي إطار يهدف إلى تعزيز الموثوقية والقابلية للتفسير في تصنيف أورام الدماغ في السيناريوهات التي قد تكون فيها بعض وسائط التصوير بالرنين المغناطيسي مفقودة. تستخدم MoRE-Net بشكل مبتكر مشفرات لكل وسيلة، مما يسمح بمعالجة كل تسلسل تصوير بالرنين المغناطيسي بشكل مستقل، مما يقلل من فقدان المعلومات ويزيد من الموثوقية. بالإضافة إلى ذلك، يتم دمج وحدة إسقاط الوسائط خلال التدريب لتسهيل تعلم تمثيلات فعالة من المدخلات غير المكتملة. للحفاظ على السياق العالمي الضروري لتصنيف مستوى الحجم، يتضمن الإطار استراتيجيتين: مشفرات قائمة على Mamba لنمذجة الاعتماديات بعيدة المدى داخل كل وسيلة، ومعلم CNN متعدد الوسائط للتقريب المبكر الذي يوجه المشفرات لكل وسيلة من خلال خسارة المحاذاة، مما يجمع بين مزايا الدمج المبكر مع تعزيز الموثوقية خلال الاستدلال.
يقيم المؤلفون MoRE-Net مقابل مجموعة من الطرق الأساسية، بما في ذلك CNN التقليدية المفسرة عبر خرائط تنشيط الفئة المعتمدة على الوزن التدرجي (Grad-CAM) ونماذج قائمة على الأجزاء القابلة للتفسير مثل ProtoPNet وXProtoNet وMProtoNet. تهدف هذه التحليل المقارن إلى تقييم شامل لكل من أداء التصنيف والقابلية للتفسير للطريقة المقترحة في ظل ظروف الوسائط المفقودة.
نتائج
في قسم النتائج، يقدم المؤلفون تحليلًا مقارنًا لقابلية تفسير نموذجهم، MoRE-Net، تحت ظروف مختلفة من فقدان الوسائط. توضح تصور خرائط النسبة أن MoRE-Net ينتج خرائط تنشيط أنظف وأكثر دقة تتماشى عن كثب مع الآفات الحقيقية عبر أربعة إعدادات وسائط. بشكل محدد، عند استخدام تركيبات مثل T1CE + FLAIR وT1CE + FLAIR + T2، يركز MoRE-Net بفعالية على النواة المعززة للآفات، بينما تظهر نماذج بديلة مثل MProtoNet أنماط تنشيط متفرقة، وتنتج نماذج أخرى، بما في ذلك XProtoNet وProtoPNet وCNN مع Grad-CAM، استجابات خشنة تمتد إلى الأنسجة الطبيعية.
تشير النتائج إلى أن MoRE-Net وMProtoNet يحققان أدق تحديدات وأقل تنشيطات زائفة عندما تكون جميع الوسائط متاحة، متطابقة عن كثب مع الحقيقة الأرضية. من الجدير بالذكر، عندما يتم استخدام T1CE فقط، يتم توجيه انتباه النموذج بشكل أساسي بواسطة إشارات التعزيز، مما يتداخل بدقة مع نواة الورم في مناطق التعزيز ولكنه يصبح أقل موثوقية في المناطق غير المعززة. يؤدي إضافة FLAIR إلى تحويل تركيز النموذج إلى وذمة شديدة الكثافة، مما يشير إلى استجابة تعويضية من خلال إشارات التسلل، بينما يساهم دمج T2 في استقرار وتعزيز تحديد الآفة. بشكل عام، تؤكد النتائج على الأهمية السريرية وموثوقية MoRE-Net في تفسير بيانات التصوير متعددة الوسائط بدقة.
مناقشة
في هذه الدراسة، يقدم المؤلفون MoRE-Net، وهو إطار موثوق وقابل للتفسير مصمم لتصنيف أورام الدماغ، خاصة في ظل ظروف فقدان وسائط التصوير بالرنين المغناطيسي. يستخدم الإطار مشفرات قائمة على Mamba لكل وسيلة تلتقط بفعالية الاعتماديات بعيدة المدى مع الحفاظ على التعقيد الخطي، مما يسمح بتعلم تمثيل موثوق من الوسائط الفردية. لتعزيز الأداء خلال التدريب المبكر، يتم دمج خسارة المحاذاة، التي تشجع على الاتساق بين التمثيلات لكل وسيلة وفرع الدمج المبكر القائم على CNN. بالإضافة إلى ذلك، يحاكي آلية إسقاط الوسائط فقدان الوسائط، مما يعزز موثوقية النموذج تجاه المدخلات غير المكتملة.
تظهر النتائج التجريبية أن MoRE-Net يتفوق بشكل كبير على الطرق الحالية في كل من دقة التشخيص والقابلية للتفسير عبر سيناريوهات فقدان الوسائط المختلفة. تشير الدراسات الاستبعادية إلى أنه بينما توفر المشفرات لكل وسيلة مرونة، فإن خسارة المحاذاة وإسقاط الوسائط أمران حاسمان لتحقيق أداء عالٍ. يتناول المؤلفون أيضًا المخاوف العملية المتعلقة بفرضيات تدريب النموذج وقابليته للتطبيق على مجموعات البيانات السريرية الواقعية، مما يظهر أن MoRE-Net يمكن أن يتكيف بفعالية مع الوسائط المفقودة. بشكل عام، تشير النتائج إلى أن MoRE-Net هو أداة واعدة للتطبيقات السريرية، حيث يقدم نتائج تشخيصية موثوقة ونتائج قابلة للتفسير حتى عندما تكون بعض الوسائط غير متاحة. ستركز الأعمال المستقبلية على توسيع تطبيقاته ودمج تقدير عدم اليقين لتعزيز موثوقية السريرية بشكل أكبر.
DOI: https://doi.org/10.2463/mrms.mp.2025-0107
PMID: https://pubmed.ncbi.nlm.nih.gov/41672496
Publication Date: 2026-01-01
Author(s): Binghua Li et al.
Primary Topic: Brain Tumor Detection and Classification
Overview
The research presents the Modality-Robust and Explainable Network (MoRE-Net), an innovative framework designed to enhance the robustness and interpretability of multimodal medical imaging diagnostic models, specifically for brain tumor grading under conditions of missing modalities. The model employs independent per-modality encoders to extract modality-specific features and incorporates two key strategies: Mamba-based encoders for efficient global-context modeling and an online multimodal teacher that utilizes an alignment loss during early training, which is gradually phased out.
Evaluated on the BraTS2020 dataset with 369 subjects, MoRE-Net achieved an average balanced accuracy (BAC) of 73.5% and activation precision (AP) of 61.2%, outperforming baseline methods by approximately 15% and 21%, respectively. Validation on the ReMIND dataset and ablation studies further confirmed the effectiveness of the proposed strategies and the overall robustness of the model. The findings indicate that MoRE-Net not only excels in diagnostic accuracy but also maintains interpretability, highlighting its potential for clinical application in high-stakes medical environments. Future research will aim to expand its applications and integrate uncertainty estimation to enhance clinical reliability.
Introduction
The introduction of the paper highlights the growing role of artificial intelligence (AI) in medical image analysis, particularly in diagnostic and prognostic applications. While AI has demonstrated significant potential, ensuring the trustworthiness of these systems is critical, especially in clinical settings. Two essential requirements for trustworthiness are interpretability—where the model’s reasoning is made explicit through tools like attribution maps and counterfactual analyses—and robustness, which refers to the model’s ability to perform reliably under various conditions, including input perturbations and distribution shifts.
The paper focuses on MRI-based 3D brain tumor analysis, which involves complex tasks that necessitate both interpretability and robustness. Although previous studies have made strides in enhancing either interpretability or robustness, there is a notable gap in research addressing both aspects simultaneously, particularly in the context of 3D brain tumor grading with missing modalities. The authors propose to improve diagnostic accuracy and interpretability while ensuring that robustness techniques do not compromise the model’s ability to capture long-range dependencies across the entire volume, which is crucial for global classification tasks.
Methods
In this study, the authors introduce the Modality-Robust and Explainable Network (MoRE-Net), a framework aimed at enhancing the robustness and interpretability of brain tumor grading in scenarios where certain MRI modalities may be missing. MoRE-Net innovatively employs per-modality encoders, allowing each MRI sequence to be processed independently, which reduces information loss and increases robustness. Additionally, a modality dropout module is integrated during training to facilitate the learning of effective representations from incomplete modality inputs. To maintain global context essential for volume-level classification, the framework incorporates two strategies: Mamba-based encoders for modeling long-range dependencies within each modality, and a multimodal early-fusion CNN teacher that guides the per-modality encoders through an alignment loss, thereby combining the advantages of early fusion with enhanced robustness during inference.
The authors evaluate MoRE-Net against a range of baseline methods, including conventional CNNs interpreted via gradient-weighted class activation mapping (Grad-CAM) and various interpretable prototypical part-based models such as ProtoPNet, XProtoNet, and MProtoNet. This comparative analysis aims to thoroughly assess both the classification performance and interpretability of the proposed method under conditions of missing modalities.
Results
In the results section, the authors present a comparative analysis of the interpretability of their model, MoRE-Net, under various missing modality conditions. The visualization of attribution maps demonstrates that MoRE-Net generates cleaner and more precise activation maps that align closely with the ground-truth lesions across four modality settings. Specifically, when using combinations such as T1CE + FLAIR and T1CE + FLAIR + T2, MoRE-Net effectively focuses on the enhancing core of lesions, while alternative models like MProtoNet exhibit scattered activation patterns, and others, including XProtoNet, ProtoPNet, and CNN with Grad-CAM, produce coarse responses that extend into normal tissue.
The findings indicate that MoRE-Net and MProtoNet yield the sharpest delineations and the fewest spurious activations when all modalities are available, closely matching the ground truth. Notably, when only T1CE is utilized, the model’s attention is primarily directed by enhancing cues, accurately overlapping with the tumor core in areas of enhancement but becoming less reliable in non-enhancing regions. The addition of FLAIR shifts the model’s focus to hyperintense edema, indicating a compensatory response through infiltration signals, while incorporating T2 further stabilizes and enhances the delineation of the lesion. Overall, the results underscore the clinical relevance and robustness of MoRE-Net in accurately interpreting multimodal imaging data.
Discussion
In this study, the authors introduce MoRE-Net, a modality-robust and interpretable framework designed for brain tumor grading, particularly under conditions of missing MRI modalities. The framework employs Mamba-based per-modality encoders that effectively capture long-range dependencies while maintaining linear complexity, allowing for robust representation learning from individual modalities. To enhance performance during early training, an alignment loss is integrated, which encourages consistency between the per-modality representations and a CNN-based early fusion branch. Additionally, a modality dropout mechanism simulates missing modalities, further improving the model’s robustness to incomplete inputs.
Experimental results demonstrate that MoRE-Net significantly outperforms existing methods in both diagnostic accuracy and interpretability across various missing modality scenarios. The ablation studies indicate that while per-modality encoders provide flexibility, the alignment loss and modality dropout are crucial for achieving high performance. The authors also address practical concerns regarding the model’s training assumptions and its applicability to real-world clinical datasets, showing that MoRE-Net can adapt to missing modalities effectively. Overall, the findings suggest that MoRE-Net is a promising tool for clinical applications, offering reliable diagnostic outputs and interpretable results even when certain modalities are unavailable. Future work will focus on expanding its applications and incorporating uncertainty estimation to enhance clinical reliability further.
