DOI: https://doi.org/10.1038/s41598-025-28813-w
PMID: https://pubmed.ncbi.nlm.nih.gov/41326486
تاريخ النشر: 2025-12-01
المؤلف: Govind Ram Chhimpa وآخرون
الموضوع الرئيسي: كشف وتصنيف أورام الدماغ
نظرة عامة
تقدم ورقة البحث طريقة متقدمة لاكتشاف أورام الدماغ باستخدام بنية YOLOv10، مع التأكيد على أهمية الأنظمة الآلية في التصوير الطبي للتشخيص المبكر وتخطيط العلاج. تتناول الدراسة قيود التحليل اليدوي، الذي غالبًا ما يكون بطيئًا وعرضة للأخطاء، من خلال استخدام نهج التعلم الانتقالي الذي يبدأ بنموذج مع معلمات من مجموعة بيانات COCO ويقوم بضبطه على مجموعة بيانات محددة لأورام الدماغ. يحقق النموذج المعدل دقة متوسطة (mAP) تبلغ 96.1% ودقة تبلغ 96.8%، متفوقًا بشكل كبير على نموذج YOLOv10 الأصلي. يبرز هذا فعالية التعلم الانتقالي في التصوير الطبي، خاصة في السيناريوهات التي تحتوي على بيانات محدودة.
في الختام، يوضح المؤلفون تنفيذ تقنية تجميد الطبقات الانتقائية لتحقيق توازن بين الميزات العامة والأنماط الخاصة بالمجال، مما يؤدي إلى تحسين قدرات الكشف، خاصة لحدود الأورام المعقدة. على الرغم من تحقيق حساسية عالية (استرجاع = 92.7%)، قد يحد التكلفة الحسابية المعتدلة للنموذج (حوالي 85.6 GFLOPs لكل صورة 640×640) من استخدامه في الأجهزة الطبية ذات الطاقة المنخفضة. ستركز الأعمال المستقبلية على تحسين النموذج للنشر السريري، ومعالجة تحديات حجم مجموعة البيانات والتباين بين المؤسسات، وتعزيز الحساسية من خلال تقنيات التصوير متعدد الأنماط والذكاء الاصطناعي القابل للتفسير. تؤكد الدراسة على إمكانية تكييف أساليب الكشف عن الكائنات المتطورة للتطبيقات السريرية، مما يمهد الطريق لأدوات تشخيصية موثوقة وقابلة للتفسير مدعومة بالذكاء الاصطناعي في الرعاية الصحية.
مقدمة
تسلط مقدمة ورقة البحث هذه الضوء على التأثير التحويلي للذكاء الاصطناعي والتعلم العميق على التصوير الطبي، خاصة في الكشف المبكر عن الآفات، وهو أمر حاسم لنتائج العلاج الفعالة. اعتمدت العملية التشخيصية تقليديًا على أطباء الأشعة المدربين، وقد تم تعزيزها بشكل كبير بواسطة الشبكات العصبية العميقة، حيث ظهرت عائلة نماذج YOLO (You Only Look Once) كإطار رائد للكشف عن الكائنات في الوقت الفعلي. تُظهر النسخة الأحدث، YOLOv10، تحسينات كبيرة في الدقة مع الحاجة إلى طاقة حسابية أقل. ومع ذلك، تواجه تطبيقات مثل هذه النماذج في المهام المتخصصة مثل اكتشاف أورام الدماغ تحديات بسبب محدودية توفر مجموعات بيانات التصوير الطبي الكبيرة والموسومة، بالإضافة إلى تنظيمات الخصوصية وتعقيد البيانات الطبية.
لمعالجة هذه التحديات، تقترح الورقة ضبط بنية YOLOv10 لاكتشاف أورام الدماغ باستخدام التعلم الانتقالي، مستفيدة من المعلمات المدربة مسبقًا من مجموعة بيانات COCO (الأشياء الشائعة في السياق). تستخدم الدراسة مجموعة بيانات أورام الدماغ Br35H، التي توفر، على الرغم من حجمها الأصغر، مجموعة مركزة من بيانات التصوير بالرنين المغناطيسي للتقييم. تشير النتائج إلى مقاييس أداء مثيرة للإعجاب، حيث تحقق 96.1% دقة متوسطة (mAP) و96.8% دقة، متجاوزة الأداء الأساسي لـ YOLOv10 على نفس مجموعة البيانات. لا يعزز هذا النهج من دقة الكشف فحسب، بل يحافظ أيضًا على قدرات الاستدلال في الوقت الفعلي، مما يجعله مناسبًا للتطبيقات السريرية. تشير النتائج إلى أن تكييف نماذج الكشف عن الكائنات العامة للمهام الطبية المتخصصة يمكن أن يحسن بشكل كبير من العمليات التشخيصية ويدعم المهنيين الطبيين في اتخاذ القرار.
طرق
في هذه الدراسة، يقترح المؤلفون منهجية لاكتشاف أورام الدماغ من خلال تكييف بنية YOLOv10 من خلال استراتيجية التعلم الانتقالي. يبدأون بنموذج YOLOv10 المدرب مسبقًا على مجموعة بيانات COCO ويستخدمون تقنية تجميد الطبقات الانتقائية، حيث يتم تجميد الطبقات الأولية المسؤولة عن اكتشاف الميزات البصرية الأساسية—مثل الحواف، والملمس، والأشكال الأساسية—أثناء التدريب. يستفيد هذا النهج من قابلية نقل الميزات منخفضة المستوى عبر مجالات مختلفة، بما في ذلك الصور الطبيعية والطبية.
بالنسبة لمجموعة البيانات، استخدم الباحثون مجموعة بيانات Br35H، التي تم تقسيمها إلى مجموعات تدريب، والتحقق، والاختبار بنسبة 70/20/10. لتعزيز قوة النموذج وتنوعه، قاموا بإدراج عينات مشروحة يدويًا من مجموعة بيانات أورام الدماغ المتاحة على كاجل، مما أدى إلى توسيع مجموعة بيانات التدريب وإدخال تنوعات إضافية في مورفولوجيا الأورام وظروف التصوير. يتم وضع علامة على كل صورة بإحداثيات صندوق محيط تحدد بدقة مناطق الأورام، مع الالتزام بتنسيق التعليق التوضيحي لـ YOLO الذي يحدد إحداثيات المركز \((x, y)\) والأبعاد (العرض، الارتفاع) لكل منطقة ورم، والتي تم تطبيعها لتناسب أبعاد الصورة.
نتائج
تظهر نتائج الدراسة فعالية وقوة نموذج YOLOv10 المقترح لاكتشاف أورام الدماغ، كما يتضح من التحليل الإحصائي ومقاييس التقييم. باستخدام نهج إعادة أخذ العينات غير المعلمي مع 1,000 تكرار، أظهر النموذج اتساقًا عاليًا عبر عدة عمليات، مما ينعكس في فترات ثقة ضيقة بنسبة 95%. تكشف منحنيات التقييم عن علاقة قوية بين الدقة والثقة، حيث تستقر عند دقة تبلغ 1.0 بعد عتبة ثقة تبلغ حوالي 0.75، مما يشير إلى معدل إيجابيات كاذبة منخفض وهو أمر حاسم للتطبيقات السريرية. تُظهر منحنى الاسترجاع والثقة أن الاسترجاع يبقى فوق 0.9 لقيم الثقة تصل إلى 0.8، مما يؤكد قدرة النموذج على اكتشاف معظم حالات الأورام الحقيقية.
حقق النموذج mAP@0.5 قدره 0.961، مما يشير إلى تحسينات متزامنة في الدقة والاسترجاع، وهو أمر مهم نظرًا لحجم مجموعة بيانات Br35H المحدود. بلغت منحنى F1-Confidence ذروته عند درجة F1 تبلغ 0.91 حول عتبة ثقة تبلغ 0.5، مما يمثل توازنًا مثاليًا بين الدقة والاسترجاع. بالإضافة إلى ذلك، أسفر منحنى ROC-AUC عن قيمة AUC تبلغ 0.98، مما يدل على قوة تمييز ممتازة بين مناطق الورم وغير الورم. أدى دمج التعلم الانتقالي إلى تحسين دقة الكشف بشكل كبير مع تقليل وقت التدريب، كما عززت تقنيات التحسين المحددة أداء النموذج. بشكل عام، تؤكد النتائج ملاءمة النموذج السريرية، مع التأكيد على دقته العالية وحساسيته، وهما أمران حاسمان للتشخيص الطبي الفعال. قد تركز الأعمال المستقبلية على تحسين هذه النتائج من خلال تقنيات تحسين متقدمة وطرق تجميع لتعزيز قوة النموذج بشكل أكبر.
مناقشة
تؤكد قسم المناقشة في ورقة البحث على الإمكانيات التحويلية للتعلم العميق في معالجة الصور الطبية، خاصة لاكتشاف أورام الدماغ. تبرز التحديات المتعلقة بدمج هذه النماذج في سير العمل السريري بسبب الطلبات الحسابية العالية وندرة مجموعات بيانات التصوير الطبي عالية الجودة. يقدم المؤلفون مساهماتهم الرئيسية، بما في ذلك ضبط نموذج YOLOv10-X على مجموعة بيانات أورام الدماغ Br35h باستخدام التعلم الانتقالي، مما يعزز الدقة مع الحفاظ على الكفاءة الحسابية. يقدمون تحليلًا مقارنًا لنماذج YOLOv10-X مع وبدون التعلم الانتقالي، مما يوضح مزايا الأوزان المدربة مسبقًا في تحقيق تقارب أسرع وتحسين التعميم، مما يجعل النموذج مناسبًا للتطبيقات الطبية في الوقت الفعلي.
تستعرض القسم أيضًا تطور معالجة الصور الطبية من الطرق التقليدية إلى أساليب التعلم العميق، مشيرة إلى قيود التقنيات السابقة التي اعتمدت بشكل كبير على استخراج الميزات اليدوية وضبط المعلمات. لقد أحدث ظهور الشبكات العصبية التلافيفية (CNNs) ثورة في هذا المجال، مما مكن من التعلم الآلي التلقائي للتمثيلات والمهام الأكثر تعقيدًا مثل التقسيم الدلالي. تناقش الورقة التقدم في بنى YOLO، وخاصة YOLOv10، التي تقدم آلية رأس كشف مزدوجة تعزز دقة الكشف وكفاءته، وهو أمر حاسم لمهام التصوير الطبي. يخلص المؤلفون إلى أن استراتيجيتهم في التعلم الانتقالي الخاصة بالمجال، التي تتضمن تجميد طبقات معينة من الهيكل العظمي لـ YOLOv10-X، تحسن بشكل كبير من أداء النموذج في اكتشاف أورام الدماغ، مما يعالج التحديات التي تطرحها محدودية توفر البيانات والحاجة إلى المعالجة في الوقت الفعلي في البيئات السريرية.
DOI: https://doi.org/10.1038/s41598-025-28813-w
PMID: https://pubmed.ncbi.nlm.nih.gov/41326486
Publication Date: 2025-12-01
Author(s): Govind Ram Chhimpa et al.
Primary Topic: Brain Tumor Detection and Classification
Overview
The research paper presents an advanced method for brain tumor detection using the YOLOv10 architecture, emphasizing the importance of automated systems in medical imaging for early diagnosis and treatment planning. The study addresses the limitations of manual analysis, which is often slow and prone to errors, by employing a transfer learning approach that initializes the model with parameters from the COCO dataset and fine-tunes it on a brain tumor-specific dataset. The fine-tuned model achieves a mean Average Precision (mAP) of 96.1% and a precision of 96.8%, significantly outperforming the original YOLOv10 model. This highlights the effectiveness of transfer learning in medical imaging, particularly in scenarios with limited data.
In the conclusion, the authors detail the implementation of selective layer-freezing to balance generalized features with domain-specific patterns, resulting in improved detection capabilities, particularly for complex tumor boundaries. Despite achieving high sensitivity (recall = 92.7%), the model’s moderate computational cost (approximately 85.6 GFLOPs per 640×640 image) may restrict its use in low-power medical devices. Future work will focus on optimizing the model for clinical deployment, addressing the challenges of dataset size and inter-institutional variability, and enhancing sensitivity through multi-modal imaging and explainable AI techniques. The study underscores the potential of adapting state-of-the-art object detection methods for clinical applications, paving the way for reliable and interpretable AI-assisted diagnostic tools in healthcare.
Introduction
The introduction of this research paper highlights the transformative impact of artificial intelligence and deep learning on medical imaging, particularly in the early detection of lesions, which is crucial for effective treatment outcomes. Traditionally reliant on trained radiologists, the diagnostic process has been significantly enhanced by deep neural networks, with the YOLO (You Only Look Once) family of models emerging as a leading framework for real-time object detection. The latest iteration, YOLOv10, demonstrates substantial improvements in accuracy while requiring less computational power. However, the application of such models to specialized tasks like brain tumor detection faces challenges due to the limited availability of large-scale, labeled medical imaging datasets, compounded by privacy regulations and the complexity of medical data.
To address these challenges, the paper proposes fine-tuning the YOLOv10 architecture for brain tumor detection using transfer learning, leveraging pre-trained parameters from the COCO (Common Objects in Context) dataset. The study utilizes the Br35H brain tumor dataset, which, despite its smaller size, provides a focused collection of MRI data for evaluation. The results indicate impressive performance metrics, achieving 96.1% mean Average Precision (mAP) and 96.8% precision, surpassing the baseline performance of YOLOv10 on the same dataset. This fine-tuning approach not only enhances detection accuracy but also preserves real-time inference capabilities, making it suitable for clinical applications. The findings suggest that adapting general-purpose object detection models for specialized medical tasks can significantly improve diagnostic processes and support medical professionals in decision-making.
Methods
In this study, the authors propose a methodology for brain tumor detection by adapting the YOLOv10 architecture through a transfer learning strategy. They begin with a YOLOv10 model pre-trained on the COCO dataset and employ a selective layer freezing technique, where the initial layers responsible for detecting fundamental visual features—such as edges, textures, and basic shapes—are frozen during training. This approach leverages the transferability of low-level features across different domains, including natural and medical images.
For the dataset, the researchers utilized the Br35H dataset, which was divided into training, validation, and testing subsets in a 70/20/10 ratio. To enhance the model’s robustness and diversity, they incorporated manually annotated samples from the Brain Tumor dataset available on Kaggle, thereby expanding the training dataset and introducing additional variations in tumor morphology and imaging conditions. Each image is annotated with bounding box coordinates that accurately localize tumor regions, adhering to the YOLO annotation format that specifies the center coordinates \((x, y)\) and dimensions (width, height) of each tumor region, normalized to the image dimensions.
Results
The results of the study demonstrate the effectiveness and robustness of the proposed YOLOv10 model for brain tumor detection, as evidenced by the statistical analysis and evaluation metrics. Utilizing a nonparametric bootstrapping approach with 1,000 resampling iterations, the model exhibited high consistency across multiple runs, reflected in narrow 95% confidence intervals. The evaluation curves reveal a strong Precision-Confidence relationship, stabilizing at a precision of 1.0 beyond a confidence threshold of approximately 0.75, indicating a low false positive rate crucial for clinical applications. The Recall-Confidence curve shows that recall remains above 0.9 for confidence values up to 0.8, affirming the model’s capability to detect most true tumor instances.
The model achieved an mAP@0.5 of 0.961, indicating simultaneous enhancements in precision and recall, particularly important given the limited size of the Br35H dataset. The F1-Confidence curve peaked at an F1-score of 0.91 around a confidence threshold of 0.5, representing an optimal balance between precision and recall. Additionally, the ROC-AUC curve yielded an AUC value of 0.98, demonstrating excellent discriminative power between tumor and non-tumor regions. The incorporation of transfer learning significantly improved detection accuracy while reducing training time, and specific augmentation techniques further enhanced model performance. Overall, the findings underscore the model’s clinical suitability, emphasizing its high precision and sensitivity, which are critical for effective medical diagnostics. Future work may focus on refining these results through advanced augmentation and ensemble methods to enhance model robustness further.
Discussion
The discussion section of the research paper emphasizes the transformative potential of deep learning in medical image processing, particularly for brain tumor detection. It highlights the challenges of integrating these models into clinical workflows due to high computational demands and the scarcity of high-quality medical imaging datasets. The authors present their key contributions, including the fine-tuning of the YOLOv10-X model on the Br35h brain tumor dataset using transfer learning, which enhances accuracy while maintaining computational efficiency. They provide a comparative analysis of YOLOv10-X models with and without transfer learning, demonstrating the advantages of pre-trained weights in achieving faster convergence and improved generalization, making the model suitable for real-time medical applications.
The section also reviews the evolution of medical image processing from traditional methods to deep learning approaches, noting the limitations of earlier techniques that relied heavily on manual feature extraction and parameter tuning. The emergence of Convolutional Neural Networks (CNNs) has revolutionized the field, enabling automated representation learning and more complex tasks such as semantic segmentation. The paper discusses the advancements in YOLO architectures, particularly YOLOv10, which introduces a dual detection head mechanism that enhances detection accuracy and efficiency, crucial for medical imaging tasks. The authors conclude that their domain-specific transfer learning strategy, which involves selectively freezing layers of the YOLOv10-X backbone, significantly improves the model’s performance in detecting brain tumors, thereby addressing the challenges posed by limited data availability and the need for real-time processing in clinical settings.
