DOI: https://doi.org/10.1038/s41598-026-48825-4
PMID: https://pubmed.ncbi.nlm.nih.gov/42129232
تاريخ النشر: 2026-05-13
المؤلف: Jinghui Chen وآخرون
الموضوع الرئيسي: كشف وتصنيف أورام الدماغ
نظرة عامة
تقدم البحث YOLO-LS، وهو إطار متقدم لاكتشاف وتقسيم أورام الدماغ في صور الرنين المغناطيسي، حيث يتناول التحديات المتعلقة بالتنوع العالي في خصائص الأورام وقيود التوصيف اليدوي. يعتمد النموذج المقترح على بنية YOLO11n-seg ويقدم ثلاثة ابتكارات رئيسية: دمج ShuffleNet V1 كهيكل خفيف الوزن لتقليل التعقيد الحسابي، وآلية DySample الديناميكية لزيادة التفاصيل الدقيقة، ووحدة C3k2-PoolingFormer لتحسين دمج الميزات والتقاط السياق.
أظهرت التقييمات التجريبية على مجموعة بيانات Figshare (3,064 صورة) ومجموعة بيانات Kaggle المستقلة (300 صورة) أن YOLO-LS حقق دقة متوسطة (mAP50) تبلغ 0.953 ± 0.011، ومعامل Dice يبلغ 0.91 ± 0.01، و95% من المسافة هاوسدورف (HD95) تبلغ 4.35 ± 0.34 مم، مما يشير إلى دقة حدودية متفوقة مقارنة بالنماذج الأساسية (p < 0.05). بالإضافة إلى ذلك، قلل النموذج من التكاليف الحسابية إلى 8.1 GFLOPs، مما يمثل انخفاضًا بنسبة 15.6% مع تحسين mAP50 بمقدار 2.9 نقطة مئوية مقارنة بالأساس. كما أن التحليلات المقارنة مع الهياكل الحديثة، بما في ذلك U-Net وSwin-UNet، أكدت فعالية YOLO-LS في تحقيق توازن بين الكفاءة الحسابية وتقسيم عالي الدقة.
مقدمة
تناقش مقدمة ورقة البحث التعقيدات المرتبطة بأورام الدماغ، التي تتميز بتقدمها السريع والتحديات العلاجية الكبيرة. تظهر الأورام الأولية، بما في ذلك الدبقيات، والورم السحائي، والأورام الغدية النخامية، تنوعًا كبيرًا في سلوكها البيولوجي وغزوتها. ومن الجدير بالذكر أن الدبقيات تقدم تشخيصًا سيئًا بشكل خاص، حيث تبلغ نسبة البقاء على قيد الحياة لمدة خمس سنوات فقط 5-10%، وطبيعتها المتسللة تعقد التدخلات الجراحية. وهذا يبرز الحاجة الملحة للكشف المبكر والتحديد الدقيق لحدود الأورام لتعزيز استراتيجيات الجراحة، وتخطيط العلاج الإشعاعي، وتقييم التنبؤ. تعاني طرق التوصيف اليدوي التقليدية من كونها كثيفة العمالة وتعاني من تباين بين المراقبين، مما يبرز ضرورة وجود أساليب أكثر كفاءة في سير العمل السريري.
لمعالجة هذه التحديات، تقدم الورقة DySample، وهو خوارزمية جديدة لزيادة العينات تهدف إلى تحسين أداء نموذج YOLO11 في مهام التنبؤ. على عكس التداخل التقليدي لأقرب الجيران، تستخدم DySample تقنية زيادة عينات ديناميكية مدركة للمحتوى تعدل مواقع العينات بناءً على ميزات الصورة المحلية، مما يعزز استعادة تفاصيل الميزات. تتضمن عملية DySample إنشاء مجموعة من نقاط العينات وإعادة عينة الميزات الديناميكية، مما يحسن من استخدام الموارد الحسابية مع الحفاظ على مخرجات عالية الدقة. تستخدم الخوارزمية طبقة تلافيفية خطية لإنشاء انزلاقات ديناميكية تعمل على تحسين مواقع العينات الأولية، مما ينتج في النهاية خريطة ميزات عالية الدقة من خلال التداخل الثنائي. تهدف هذه الطريقة المبتكرة إلى تحسين دقة وكفاءة تحديد حدود الأورام في التطبيقات السريرية.
طرق
تم إجراء الطرق التجريبية الموضحة في هذه الدراسة على نظام تشغيل Windows 11 باستخدام معالج Intel(R) Core(TM) Ultra 9 275 HX، و بطاقة الرسوميات NVIDIA GeForce GTX 5070 Laptop، و 16 جيجابايت من الذاكرة العشوائية. تم استخدام إطار التعلم العميق PyTorch-2.8.0، مع CUDA-12.9 و Python-3.9.23. يتم تفصيل معلمات النموذج في الجدول 1.
لتقييم أداء النموذج في التعرف على صور الرنين المغناطيسي لأورام الدماغ وتقسيمها، تم استخدام مجموعة متنوعة من المقاييس، بما في ذلك الدقة، والاسترجاع، ومتوسط الدقة (mAP)، والتقاطع على الاتحاد (IoU)، ومعامل Dice، و95% من المسافة هاوسدورف (HD95). تم استخدام mAP50 كمقياس رئيسي، يمثل الدقة المتوسطة عند IoU=0.5، بينما قدم mAP50-95 تقييمًا أوسع عبر قيم IoU من 0.5 إلى 0.95. هذه المقاييس حاسمة لتقييم دقة النموذج في اكتشاف مواقع الأورام، وتصنيف أنواع الأورام، وتقسيم حدود الأورام، بالإضافة إلى تقييم الكفاءة الحسابية من خلال GFLOPs. يتم تقديم صيغ حساب هذه المقاييس في الجدول 2، حيث يتم تعريف HD95 كـ \( d_{Hausdorff}(X,Y) = \max \{ \max_{x \in X} \min_{y \in Y} d(x,y), \max_{y \in Y} \min_{x \in X} d(x,y) \} \). بالإضافة إلى ذلك، تم حساب متوسط IoU (mIoU) ومعامل Dice كمقاييس تكميلية، بينما تم تحليل GFLOPs لتقييم كفاءة خفة النموذج، مما يضمن قابليته للتطبيق في البيئات السريرية ذات الموارد المحدودة.
مناقشة
في قسم المناقشة من ورقة البحث، يبرز المؤلفون التقدم في التصوير بالرنين المغناطيسي (MRI) لتشخيص أورام الدماغ، مشددين على الدور الكبير للتعلم العميق، وخاصة الشبكات العصبية التلافيفية (CNNs) وهياكل المحولات، في أتمتة تحليل الرنين المغناطيسي. بينما أظهرت الشبكات العصبية التلافيفية، مثل متغيرات U-Net، وعدًا في استخراج الميزات المحلية، إلا أنها مقيدة بمجالات استقبال ثابتة، مما يحد من قدرتها على التقاط الاعتمادات بعيدة المدى. ظهرت طرق قائمة على المحولات لمعالجة هذه القيود، محققة تقدمًا في نمذجة السياق العالمي ولكن تواجه تحديات تتعلق بالتعقيد الحسابي. يشير المؤلفون إلى تطور سلسلة YOLO (You Only Look Once)، التي تم تكييفها لتقسيم أورام الدماغ في الوقت الحقيقي، ومع ذلك، تظهر المتغيرات الحالية من YOLO قيودًا في الحفاظ على التفاصيل الدقيقة، والنشر الخفيف، ودمج الميزات عبر المقاييس.
لمعالجة هذه التحديات، تقترح الدراسة YOLO-LS (تقسيم خفيف الوزن)، الذي يتضمن عدة ابتكارات: هيكل خفيف الوزن باستخدام ShuffleNet V1 لتقليل الحمل الحسابي، وآلية زيادة العينات الديناميكية (DySample) لتعزيز استعادة التفاصيل في الآفات الصغيرة، ووحدة C3k2-PoolingFormer لدمج الميزات بكفاءة. يتم التحقق من النموذج المقترح ضد كل من مجموعات البيانات الداخلية والخارجية، مما يظهر تحسينات كبيرة في مقاييس التقسيم مثل متوسط الدقة (mAP50)، ودرجة Dice، والمسافة هاوسدورف (HD95)، بينما يحقق أيضًا تقليلًا في التعقيد الحسابي. تشير النتائج إلى أن YOLO-LS لا يعزز فقط دقة التقسيم ودقة الحدود، بل يحافظ أيضًا على كفاءة مناسبة للتطبيقات السريرية، مما يعالج الحاجة الملحة لحلول معمارية متقدمة في التصوير الطبي.
القيود
تسلط القيود في الدراسة الضوء على عدة مجالات حرجة للتحسين والبحث المستقبلي. أولاً، قد تعيق الاعتماد على تسلسلات الرنين المغناطيسي المعزز بتباين T1 (CE-MRI) التحديد الدقيق للإدماج المحيط بالورم، الذي يتم تصويره بشكل أفضل باستخدام تسلسلات T2 أو FLAIR. بالإضافة إلى ذلك، بينما أظهر النموذج قدرات تعميم، فإن الحجم الصغير لمجموعة الاختبار الخارجية المستقلة يثير القلق بشأن المتانة عبر مؤسسات وبروتوكولات تصوير مختلفة. من الجدير بالذكر أن الدبقيات أظهرت استرجاعًا منخفضًا في التحقق الخارجي، يُعزى إلى طبيعتها المتسللة وحدودها غير الواضحة، مما يعقد جهود التقسيم. علاوة على ذلك، على الرغم من أن نموذج YOLO-LS يظهر انخفاضًا كبيرًا في التكلفة الحسابية (8.1 GFLOPs)، إلا أن كفاءته تم تقييمها فقط في بيئة GPU محكومة، دون إجراء نشر عملي أو تقييم على منصات الأجهزة الطرفية حتى الآن.
سيركز العمل المستقبلي على عدة اتجاهات رئيسية: دمج بيانات التصوير متعددة الأنماط (بما في ذلك T1، T2، FLAIR، DWI، وDCE-MRI) لتعزيز العلامات البيولوجية التصويرية؛ تطوير استراتيجيات تعلم تكيفية شخصية لتعديل النموذج بسرعة بناءً على التصوير الخاص بالمريض؛ ودمج نتائج التقسيم مع أنظمة تخطيط العلاج الإشعاعي (RTPS) لتحديد حجم الورم تلقائيًا. بالإضافة إلى ذلك، هناك حاجة إلى تجارب سريرية صارمة لتقييم تأثير YOLO-LS على كفاءة التشخيص ونتائج المرضى. لمعالجة مخاوف خصوصية البيانات، تخطط الدراسة لتنفيذ تقنيات التعلم الفيدرالي للتدريب التعاوني عبر المؤسسات. أخيرًا، سيتم بذل الجهود لتصدير نموذج YOLO-LS إلى تنسيقات محسّنة للتقييم على الأجهزة الطرفية، مما يضمن النشر العملي في البيئات السريرية.
DOI: https://doi.org/10.1038/s41598-026-48825-4
PMID: https://pubmed.ncbi.nlm.nih.gov/42129232
Publication Date: 2026-05-13
Author(s): Jinghui Chen et al.
Primary Topic: Brain Tumor Detection and Classification
Overview
The research presents YOLO-LS, an advanced framework for the detection and segmentation of brain tumors in MRI scans, addressing the challenges of high heterogeneity in tumor characteristics and the limitations of manual annotation. The proposed model builds on the YOLO11n-seg architecture and introduces three key innovations: the integration of ShuffleNet V1 as a lightweight backbone to lower computational complexity, the DySample dynamic upsampling mechanism to enhance fine-grained detail recovery, and the C3k2-PoolingFormer module for improved feature fusion and context capture.
Empirical evaluations on the Figshare dataset (3,064 images) and an independent Kaggle dataset (300 images) demonstrated that YOLO-LS achieved a mean Average Precision (mAP50) of 0.953 ± 0.011, a Dice coefficient of 0.91 ± 0.01, and a 95% Hausdorff Distance (HD95) of 4.35 ± 0.34 mm, indicating superior boundary accuracy compared to baseline models (p < 0.05). Additionally, the model reduced computational costs to 8.1 GFLOPs, representing a 15.6% decrease while improving mAP50 by 2.9 percentage points over the baseline. Comparative analyses with state-of-the-art architectures, including U-Net and Swin-UNet, further validated the efficacy of YOLO-LS in achieving a balance between computational efficiency and high-precision segmentation.
Introduction
The introduction of the research paper discusses the complexities associated with brain tumors, which are characterized by their rapid progression and significant therapeutic challenges. Primary tumors, including gliomas, meningiomas, and pituitary adenomas, exhibit considerable heterogeneity in their biological behavior and invasiveness. Notably, gliomas present a particularly poor prognosis, with a five-year survival rate of only 5-10%, and their infiltrative nature complicates surgical interventions. This underscores the critical need for early detection and accurate delineation of tumor boundaries to enhance surgical strategies, radiotherapy planning, and prognosis assessment. Traditional manual annotation methods are labor-intensive and suffer from inter-observer variability, highlighting the necessity for more efficient approaches in clinical workflows.
To address these challenges, the paper introduces DySample, a novel upsampling algorithm designed to improve the performance of the YOLO11 model in prediction tasks. Unlike traditional nearest-neighbor interpolation, DySample employs a content-aware dynamic upsampling technique that adjusts sampling positions based on local image features, thereby enhancing the restoration of feature details. The DySample process involves generating a set of sampling points and dynamic feature resampling, which optimizes computational resources while maintaining high-resolution outputs. The algorithm utilizes a linear convolutional layer to create dynamic offsets that refine the initial sampling positions, ultimately producing a high-resolution feature map through bilinear interpolation. This innovative approach aims to improve the accuracy and efficiency of tumor boundary delineation in clinical applications.
Methods
The experimental methods outlined in this study were conducted on a Windows 11 operating system utilizing an Intel(R) Core(TM) Ultra 9 275 HX CPU, an NVIDIA GeForce GTX 5070 Laptop GPU, and 16 GB of RAM. The deep learning framework employed was PyTorch-2.8.0, with CUDA-12.9 and Python-3.9.23. Model hyperparameters are detailed in Table 1.
To evaluate the performance of the model in brain tumor MRI image recognition and segmentation, a variety of metrics were utilized, including precision, recall, mean average precision (mAP), intersection over union (IoU), Dice Coefficient, and 95% Hausdorff Distance (HD95). Specifically, mAP50 was used as the primary metric, representing average precision at IoU=0.5, while mAP50-95 provided a broader evaluation across IoU values from 0.5 to 0.95. These metrics are critical for assessing the model’s accuracy in detecting tumor locations, classifying tumor types, and segmenting tumor boundaries, as well as evaluating computational efficiency through GFLOPs. The calculation formulas for these metrics are presented in Table 2, with HD95 defined as \( d_{Hausdorff}(X,Y) = \max \{ \max_{x \in X} \min_{y \in Y} d(x,y), \max_{y \in Y} \min_{x \in X} d(x,y) \} \). Additionally, mean IoU (mIoU) and the Dice coefficient were computed as supplementary metrics, while GFLOPs were analyzed to assess the model’s lightweighting efficiency, ensuring its applicability in resource-constrained clinical settings.
Discussion
In the discussion section of the research paper, the authors highlight the advancements in Magnetic Resonance Imaging (MRI) for brain tumor diagnosis, emphasizing the significant role of deep learning, particularly Convolutional Neural Networks (CNNs) and Transformer architectures, in automating MRI analysis. While CNNs, such as U-Net variants, have shown promise in local feature extraction, they are constrained by fixed receptive fields, limiting their ability to capture long-range dependencies. Transformer-based methods have emerged to address these limitations, achieving breakthroughs in global context modeling but facing challenges related to computational complexity. The authors note the evolution of the YOLO (You Only Look Once) series, which has been adapted for real-time brain tumor segmentation, yet existing YOLO variants exhibit limitations in preserving fine details, lightweight deployment, and cross-scale feature fusion.
To address these challenges, the study proposes YOLO-LS (Lightweight Segmentation), which incorporates several innovations: a lightweight backbone using ShuffleNet V1 to reduce computational load, a dynamic upsampling mechanism (DySample) to enhance detail recovery in small lesions, and a C3k2-PoolingFormer module for efficient feature fusion. The proposed model is validated against both internal and external datasets, demonstrating significant improvements in segmentation metrics such as mean Average Precision (mAP50), Dice score, and Hausdorff Distance (HD95), while also achieving a reduction in computational complexity. The results indicate that YOLO-LS not only enhances segmentation precision and boundary accuracy but also maintains efficiency suitable for clinical applications, thereby addressing the pressing need for advanced architectural solutions in medical imaging.
Limitations
The limitations of the study highlight several critical areas for improvement and future research. Firstly, the reliance on T1-weighted contrast-enhanced MRI (CE-MRI) sequences may hinder the accurate delineation of peritumoral edema, which is better visualized with T2-weighted or FLAIR sequences. Additionally, while the model demonstrated generalization capabilities, the small size of the external independent test set raises concerns about robustness across various institutions and imaging protocols. Notably, gliomas exhibited a reduced recall in external validation, attributed to their infiltrative nature and indistinct boundaries, complicating segmentation efforts. Furthermore, although the YOLO-LS model shows a significant reduction in computational cost (8.1 GFLOPs), its efficiency has only been evaluated in a controlled GPU environment, with no practical deployment or benchmarking on edge hardware platforms conducted yet.
Future work will focus on several key directions: integrating multimodal imaging data (including T1, T2, FLAIR, DWI, and DCE-MRI) to enhance imaging biomarkers; developing personalized adaptive learning strategies for rapid model fine-tuning based on patient-specific imaging; and integrating segmentation results with Radiation Therapy Planning Systems (RTPS) for automatic tumor volume delineation. Additionally, rigorous clinical trials are needed to assess the impact of YOLO-LS on diagnostic efficiency and patient outcomes. To address data privacy concerns, the study plans to implement Federated Learning techniques for collaborative training across institutions. Lastly, efforts will be made to export the YOLO-LS model to optimized formats for benchmarking on edge devices, ensuring practical deployment in clinical settings.
