DOI: https://doi.org/10.1016/j.patcog.2026.113647
تاريخ النشر: 2026-04-04
المؤلف: Yuan Ma وآخرون
الموضوع الرئيسي: تعلم الآلة وتصنيف البيانات
نظرة عامة
يقدم هذا القسم نظرة عامة على التحديات المرتبطة بالتعلم من التسميات الضوضائية في تحليل الصور الطبية، حيث المعرفة الخبيرة مطلوبة للتعليق وغالبًا ما تؤدي التباينات بين المراقبين إلى تسميات غير متسقة. لتقييم قوة الطرق الحالية في هذا المجال بشكل منهجي، يقدم المؤلفون LNMBench، وهو معيار شامل مصمم للضوضاء في التسميات في التصوير الطبي. يتضمن LNMBench 10 طرق تمثيلية تم اختبارها عبر 7 مجموعات بيانات، و6 أوضاع تصوير، و3 أنماط ضوضاء، مما يوفر إطارًا موحدًا لتقييم القوة تحت ظروف واقعية. تشير النتائج التجريبية إلى أن طرق التعلم الحالية مع التسميات الضوضائية (LNL) تؤدي بشكل ملحوظ في ظل وجود ضوضاء عالية وعالمية، مما يزيد من تفاقم مشاكل عدم توازن الفئات وتباين المجال.
استجابةً لهذه النتائج، يقترح المؤلفون تحسينًا بسيطًا ولكنه فعال لتحسين قوة النموذج ضد هذه الضوضاء. يؤكدون على أهمية قاعدة الشيفرة المتاحة للجمهور LNMBench، التي تهدف إلى تسهيل التقييمات القياسية وتعزيز البحث القابل للتكرار. تعيد الخاتمة التأكيد على قيود طرق LNL الحالية في السيناريوهات الطبية الواقعية وتبرز تقديم طريقة شبه إشرافية مصممة لتقليل تلوث المجموعة النظيفة وتعزيز الأداء للفئات الأقل في التصوير الطبي. يهدف هذا العمل إلى دعم مجتمع البحث وإلهام التقدم في تقنيات التعلم القوي للتسميات الضوضائية في التطبيقات الطبية.
مقدمة
تسلط مقدمة هذه الورقة البحثية الضوء على أهمية التعلم العميق في رؤية الكمبيوتر، خاصة في سياق التصوير الطبي، حيث تعتبر ضوضاء التسميات تحديًا شائعًا بسبب الحاجة إلى التعليقات الخبيرة. يناقش المؤلفون طرقًا مختلفة تم تطويرها لمعالجة ضوضاء التسميات، بما في ذلك نمذجة تدهور التسميات من خلال مصفوفات انتقال الضوضاء، ودوال الخسارة القوية، واستراتيجيات اختيار العينات بناءً على معيار الخسارة الصغيرة. بينما أظهرت هذه الأساليب وعدًا في الصور الطبيعية، لا تزال فعاليتها في التصوير الطبي غير مستكشفة بشكل كافٍ.
لسد هذه الفجوة، يقدم المؤلفون LNMBench، وهو معيار شامل مصمم لتقييم عشرة طرق تمثيلية لتعلم ضوضاء التسميات (LNL) عبر سبع مجموعات بيانات للتصوير الطبي. يصنف هذا المعيار ضوضاء التسميات إلى مصادر اصطناعية وعالمية ويقيم أداء النموذج تحت ظروف متنوعة، بما في ذلك عدم توازن الفئات النموذجي في مجموعات البيانات الطبية. تكشف النتائج أن طرق LNL الحالية غير كافية لإدارة ضوضاء التسميات، خاصة في السيناريوهات ذات الضوضاء العالية، وتؤكد تأثير عدم توازن الفئات على الأداء. يقترح المؤلفون MedSSL، وهي استراتيجية جديدة لتعزيز الطرق شبه الإشرافية، مما يظهر تحسينًا في الأداء على مجموعات البيانات الضوضائية الواقعية. تشمل مساهمات هذا العمل إنشاء إطار تقييم موحد، وتصنيف طرق LNL، وقاعدة شيفرة سهلة الاستخدام لتسهيل المزيد من البحث في هذا المجال الحرج.
طرق
في هذا القسم، يقدم المؤلفون تقييمًا منهجيًا لمختلف طرق التصنيف تحت ثلاثة أنماط ضوضاء متميزة: الضوضاء المتماثلة، الضوضاء المعتمدة على الحالة، والضوضاء الواقعية. تم حساب متوسط دقة التصنيف على مدى الخمس دورات الأخيرة لكل طريقة، مما يسمح بمقارنة مستقرة وعادلة. تم تصنيف النتائج، وحساب متوسط الدقة عبر جميع أنماط الضوضاء لتقييم القوة العامة لكل طريقة.
تشير النتائج إلى أن Co-teaching وDISC حققت باستمرار أعلى متوسط دقة تصنيف عبر جميع ظروف الضوضاء. على وجه الخصوص، تفوقت Co-teaching+ وDivideMix تحت الضوضاء المتماثلة، بينما كانت Co-teaching وCoDis الأفضل تحت الضوضاء المعتمدة على الحالة. من الجدير بالذكر أن طرق التنظيم القوية ضد الضوضاء، SCE وCDR، أظهرت أداءً متفوقًا في ظل وجود الضوضاء الواقعية. تبرز هذه النتائج فعالية بعض الطرق في التخفيف من تأثير أنواع الضوضاء المختلفة على دقة التصنيف.
نتائج
في هذا القسم، يتم تقييم أداء مختلف طرق تعلم ضوضاء التسميات (LNL) على مجموعة بيانات PathMNIST المتوازنة تحت ظروف ضوضاء اصطناعية، تحديدًا الضوضاء المتماثلة والضوضاء المعتمدة على الحالة. تشير النتائج إلى أنه بينما تؤدي جميع الطرق بشكل جيد تحت ضوضاء متماثلة منخفضة (Sym-20%)، فإن الأداء ينخفض بشكل ملحوظ مع زيادة الضوضاء، خاصة عند Sym-50% وSym-90%، حيث تكافح معظم الطرق، خاصة تلك التي تعتمد على مصفوفات الانتقال وتقنيات التنظيم. من الجدير بالذكر أن طرقًا مثل Co-teaching وCoDis والطرق شبه الإشرافية تحافظ على أداء أفضل تحت مستويات الضوضاء العالية، على الرغم من أن جميع الطرق تنهار في النهاية عند Sym-90%. في سيناريو الضوضاء المعتمدة على الحالة، تحتفظ معظم الطرق بدقة معقولة عند Idn-20%، ولكن تواجه انخفاضات كبيرة في الأداء عند Idn-50% وفشل كامل عند Idn-90%، مما يبرز الفعالية المحدودة لطرق التنظيم ومصفوفات الانتقال.
يمتد التحليل إلى مجموعات البيانات غير المتوازنة، حيث يتم تقييم طرق LNL على مجموعات بيانات التصوير الطبي (DermaMNIST وBloodMNIST وOrganCMNIST) تحت الضوضاء المعتمدة على الحالة. تظهر النتائج أنه بينما يتم تحقيق دقة عالية تحت Idn-20% لـ BloodMNIST، فإن فقط Co-teaching وCoDis والطرق شبه الإشرافية تحافظ على الأداء عند Idn-50%. تحت Idn-90%، تظهر جميع الطرق تدهورًا كبيرًا. بالنسبة لـ DermaMNIST، تؤدي CDR بشكل جيد في البداية عند Idn-20% ولكن تعاني عند Idn-50%، بينما تظهر DISC وCoDis وCo-teaching اتجاهًا مشابهًا. تشير النتائج إلى أن طرق اختيار العينات أقل فعالية في ظل وجود ضوضاء عالية، خاصة في الإعدادات غير المتوازنة، كما يتضح من أدائها بالنسبة لقاعدة الانتروبيا المتقاطعة القياسية (CE). أخيرًا، ينتهي القسم بتقييم طرق LNL على مجموعات البيانات الضوضائية الواقعية، مما يشير إلى الآثار العملية للنتائج.
مناقشة
يوفر قسم المناقشة في الورقة نظرة شاملة على المنهجيات الحالية لتعلم التسميات الضوضائية، خاصة في سياق تحليل الصور الطبية. يصنف هذه الطرق إلى أربعة نهج رئيسية: تقدير مصفوفة انتقال الضوضاء، التنظيم القوي ضد الضوضاء، اختيار العينات، والطرق شبه الإشرافية. من الجدير بالذكر أن تقنيات تقدير مصفوفة انتقال الضوضاء، مثل T-Revision، تهدف إلى استنتاج توزيعات التسميات النظيفة ولكن غالبًا ما تكافح في السيناريوهات الواقعية بسبب غياب نقاط مرجعية موثوقة. ركزت التطورات الأخيرة على مصفوفات الانتقال المعتمدة على الحالة، التي تعزز القوة من خلال السماح بنمذجة الضوضاء لتختلف عبر العينات. تعد طرق التنظيم القوية ضد الضوضاء، مثل SCE، تعدل دوال الخسارة للتخفيف من الإفراط في التكيف مع التسميات الضوضائية، بينما تحدد استراتيجيات اختيار العينات، التي تمثلها Co-teaching، وتستخدم العينات النظيفة المحتملة لتحسين أداء النموذج.
يسلط القسم الضوء أيضًا على التحديات التي تواجه تحليل الصور الطبية، حيث تعتبر ضوضاء التسميات شائعة بسبب تعقيدات التعليق البشري. تم اقتراح أطر مختلفة لمعالجة هذه القضايا، بما في ذلك الأساليب المدفوعة بالشكوك واستراتيجيات إزالة الضوضاء متعددة المراحل. ومع ذلك، فإن العديد من هذه الطرق محددة بمجموعات البيانات وتفتقر إلى تنفيذات موحدة، مما يعقد المقارنات العادلة. يتم التأكيد على الحاجة إلى معايير مخصصة في التصوير الطبي، حيث غالبًا ما تستمد الأساليب الحالية من مجموعات بيانات الصور الطبيعية، مما يترك فجوة في فهم فعاليتها في السياقات الطبية. يقترح المؤلفون معيارًا موحدًا لتقييم طرق تعلم ضوضاء التسميات بشكل منهجي عبر سيناريوهات التصوير الطبي المتنوعة، بهدف تحديد التحديات الرئيسية وإبلاغ اتجاهات البحث المستقبلية.
القيود
تقدم الدراسة عدة قيود تؤثر على نتائجها. من الجدير بالذكر أن غياب تنفيذات متاحة للجمهور حال دون تضمين طرق تعلم ضوضاء التسميات (LNL) المصممة لمجموعات البيانات الطبية في المعيار. بالإضافة إلى ذلك، أدت التعديلات المقترحة إلى تحسينات هامشية فقط على الطرق شبه الإشرافية الحالية، مما يشير إلى قيود محتملة في فعالية النهج الحالي.
مع النظر إلى المستقبل، ستركز الأبحاث المستقبلية على توسيع المعيار ليشمل مجموعات بيانات طبية ضوضائية واقعية وطرق LNL أكثر تخصصًا لصور الطب. علاوة على ذلك، يعتزم المؤلفون تنويع مقاييس التقييم لتعزيز قوة النتائج. كما يخططون للحفاظ على LNMBench كمنصة مفتوحة ومتطورة، مما يعزز البحث العادل والقابل للتكرار ضمن مجال تعلم ضوضاء التسميات في السياقات الطبية.
DOI: https://doi.org/10.1016/j.patcog.2026.113647
Publication Date: 2026-04-04
Author(s): Yuan Ma et al.
Primary Topic: Machine Learning and Data Classification
Overview
The section presents an overview of the challenges associated with learning from noisy labels in medical image analysis, where expert knowledge is required for annotation and inter-observer variability often results in inconsistent labels. To systematically evaluate the robustness of existing methods in this domain, the authors introduce LNMBench, a comprehensive benchmark designed for Label Noise in Medical imaging. LNMBench includes 10 representative methods tested across 7 datasets, 6 imaging modalities, and 3 noise patterns, providing a unified framework for assessing robustness under realistic conditions. Experimental results indicate that current learning with noisy labels (LNL) methods significantly underperform in the presence of high and real-world noise, exacerbated by issues of class imbalance and domain variability.
In response to these findings, the authors propose a simple yet effective enhancement to improve model robustness against such noise. They emphasize the importance of their publicly available LNMBench codebase, which aims to facilitate standardized evaluations and promote reproducible research. The conclusion reiterates the limitations of existing LNL approaches in real-world medical scenarios and highlights the introduction of a semi-supervised method designed to reduce clean set contamination and enhance performance for minority classes in medical imaging. This work aims to support the research community and inspire advancements in robust learning techniques for noisy labels in medical applications.
Introduction
The introduction of this research paper highlights the significance of deep learning in computer vision, particularly in the context of medical imaging, where label noise is a prevalent challenge due to the need for expert annotations. The authors discuss various methods developed to address label noise, including the modeling of label corruption through noise transition matrices, robust loss functions, and sample selection strategies based on the small-loss criterion. While these approaches have shown promise in natural images, their effectiveness in medical imaging remains underexplored.
To bridge this gap, the authors introduce LNMBench, a comprehensive benchmark designed to systematically evaluate ten representative label noise learning (LNL) methods across seven medical imaging datasets. This benchmark categorizes label noise into synthetic and real-world sources and assesses model performance under diverse conditions, including class imbalances typical in medical datasets. The findings reveal that existing LNL methods are insufficient for managing label noise, particularly in high-noise scenarios, and underscore the impact of class imbalance on performance. The authors propose MedSSL, a novel strategy to enhance semisupervised methods, demonstrating improved performance on real-world noisy datasets. The contributions of this work include the establishment of a standardized evaluation framework, a ranking of LNL methods, and a user-friendly codebase to facilitate further research in this critical area.
Methods
In this section, the authors present a systematic evaluation of various classification methods under three distinct noise patterns: symmetric noise, instance-dependent noise, and real-world noise. The average classification accuracy was computed over the last five epochs for each method, allowing for a stable and fair comparison. The results were ranked, and a mean accuracy across all noise patterns was calculated to assess the overall robustness of each method.
The findings indicate that Co-teaching and DISC consistently achieved the highest average classification accuracy across all noise conditions. Specifically, Co-teaching+ and DivideMix excelled under symmetric noise, while Co-teaching and CoDis performed best under instance-dependent noise. Notably, the noise-robust regularization methods, SCE and CDR, demonstrated superior performance in the presence of real-world noise. These results highlight the effectiveness of certain methods in mitigating the impact of different noise types on classification accuracy.
Results
In this section, the performance of various Label Noise Learning (LNL) methods is evaluated on the balanced PathMNIST dataset under synthetic noise conditions, specifically symmetric and instance-dependent noise. The results indicate that while all methods perform well under low symmetric noise (Sym-20%), performance declines significantly as noise increases, particularly at Sym-50% and Sym-90%, where most methods struggle, especially those relying on transition matrices and regularization techniques. Notably, methods such as Co-teaching, CoDis, and semi-supervised approaches maintain better performance under higher noise levels, although all methods ultimately collapse at Sym-90%. In the instance-dependent noise scenario, most methods retain reasonable accuracy at Idn-20%, but face substantial performance drops at Idn-50% and complete failure at Idn-90%, highlighting the limited effectiveness of regularization and transition matrix methods.
The analysis extends to imbalanced datasets, where LNL methods are assessed on medical imaging datasets (DermaMNIST, BloodMNIST, and OrganCMNIST) under instance-dependent noise. Results show that while high accuracy is achieved under Idn-20% for BloodMNIST, only Co-teaching, CoDis, and semi-supervised methods maintain performance at Idn-50%. Under Idn-90%, all methods exhibit significant degradation. For DermaMNIST, CDR initially performs well at Idn-20% but suffers at Idn-50%, while DISC, CoDis, and Co-teaching show a similar trend. The findings suggest that sample selection methods are less effective in the presence of high noise, particularly in imbalanced settings, as evidenced by their performance relative to the standard cross-entropy (CE) baseline. Finally, the section concludes with an evaluation of LNL methods on real-world noisy datasets, indicating the practical implications of the findings.
Discussion
The discussion section of the paper provides a comprehensive overview of existing methodologies for learning with noisy labels, particularly in the context of medical image analysis. It categorizes these methods into four main approaches: noise transition matrix estimation, noise-robust regularization, sample selection, and semi-supervised methods. Notably, noise transition matrix estimation techniques, such as T-Revision, aim to infer clean label distributions but often struggle in real-world scenarios due to the absence of reliable anchor points. Recent advancements have focused on instance-dependent transition matrices, which enhance robustness by allowing noise modeling to vary across samples. Noise-robust regularization methods, like SCE, modify loss functions to mitigate overfitting to noisy labels, while sample selection strategies, exemplified by Co-teaching, identify and utilize likely clean samples to improve model performance.
The section also highlights the challenges faced in medical image analysis, where label noise is prevalent due to the complexities of human annotation. Various frameworks have been proposed to address these issues, including uncertainty-driven approaches and multi-stage noise removal strategies. However, many of these methods are dataset-specific and lack standardized implementations, complicating fair comparisons. The need for dedicated benchmarks in medical imaging is emphasized, as existing approaches often derive from natural image datasets, leaving a gap in understanding their effectiveness in medical contexts. The authors propose a standardized benchmark to systematically evaluate label noise learning methods across diverse medical imaging scenarios, aiming to identify key challenges and inform future research directions.
Limitations
The study presents several limitations that impact its findings. Notably, the absence of publicly available implementations prevented the inclusion of label-noise learning (LNL) methods tailored for medical datasets in the benchmark. Additionally, the proposed modifications yielded only marginal improvements over existing semi-supervised methods, indicating potential constraints in the effectiveness of the current approach.
Looking ahead, future research will focus on expanding the benchmark to incorporate real-world noisy medical datasets and more specialized LNL methods for medical images. Furthermore, the authors intend to diversify evaluation metrics to enhance the robustness of the findings. They also plan to maintain LNMBench as an open and evolving platform, promoting fair and reproducible research within the domain of label-noise learning in medical contexts.
