DOI: https://doi.org/10.1109/tmi.2026.3673118
PMID: https://pubmed.ncbi.nlm.nih.gov/41811722
تاريخ النشر: 2026-03-11
المؤلف: Theodore Barfoot وآخرون
الموضوع الرئيسي: تطبيقات الشبكات العصبية المتقدمة
نظرة عامة
في هذا البحث، يتناول المؤلفون قضية الثقة المفرطة في الشبكات العصبية العميقة المستخدمة في تقسيم الصور الطبية، والتي يمكن أن تقوض موثوقيتها في البيئات السريرية. يقترحون استخدام صيغ قابلة للتفريق لخطأ المعايرة المتوسطة L1 (mL1-ACE) كدالة خسارة مساعدة، يتم حسابها على أساس كل صورة، لتعزيز المعايرة على مستوى البكسل. تكشف تجاربهم عبر أربعة مجموعات بيانات (ACDC، AMOS، KiTS، BraTS) أن دمج mL1-ACE يقلل بشكل فعال من كل من خطأ المعايرة المتوسطة (ACE) وخطأ المعايرة الأقصى (MCE)، مع الحفاظ على قيم عالية لمعاملات تشابه دايس (DSCs). ومن الجدير بالذكر أن النسخة ذات الصناديق الناعمة من mL1-ACE تحقق أكبر تحسينات في المعايرة، على الرغم من وجود تنازل طفيف في أداء التقسيم، بينما تحافظ النسخة ذات الصناديق الصلبة على دقة التقسيم مع تحسينات معتدلة في المعايرة.
تخلص الدراسة إلى أن دمج خسائر المعايرة أثناء التدريب يؤدي إلى تقليل كبير في أخطاء المعايرة، حيث تقلل النسخة ذات الصناديق الناعمة من mL1-ACE من ACE المتوسط والميكروي بنسبة 33% و51%، على التوالي، بتكلفة بسيطة لدقة التقسيم. بالمقابل، تحسن النسخة ذات الصناديق الصلبة المعايرة بنسبة 16% و32% لـ ACE المتوسط والميكروي، على التوالي، مع تحسين طفيف في دقة التقسيم. يؤكد المؤلفون على أهمية تحقيق التوازن بين المعايرة والدقة، مما يوفر للممارسين خيارات مصممة لتلبية الاحتياجات السريرية المحددة. كما يدعون لاستخدام خطأ المعايرة المتوسطة (ACE) بدلاً من خطأ المعايرة المتوقع (ECE) لتقييم أفضل للخطأ في المعايرة، خاصة في المناطق الحرجة. بالإضافة إلى ذلك، يقدمون مخططات موثوقية مجموعة البيانات لتصور أداء المعايرة عبر مجموعات البيانات ويسلطون الضوء على التحديات المتعلقة بفقدان فئات الحقيقة الأرضية، مقترحين طرقًا للبحث المستقبلي. بشكل عام، تؤكد نتائجهم على إمكانية تحسين المعايرة أثناء التدريب لتعزيز موثوقية أنظمة الذكاء الاصطناعي في التصوير الطبي.
مقدمة
في المقدمة، يتناول المؤلفون القضية الحرجة للخطأ في المعايرة في الشبكات العصبية العميقة (DNNs)، خاصة في سياق التصوير الطبي، حيث يمكن أن تؤدي الثقة المفرطة في التنبؤات إلى مخاطر كبيرة في اتخاذ القرارات السريرية. يؤكدون على أهمية النماذج المعايرة بشكل جيد، والتي تتماشى مع الاحتمالات المتوقعة مع النتائج الفعلية، مما يعزز موثوقية تطبيقات الذكاء الاصطناعي في الرعاية الصحية. تُلاحظ طرق المعايرة التقليدية بعد الحدث، مثل مقياس بلات وقياس الحرارة، لقيودها، حيث إنها لا تستفيد بالكامل من قدرة تعلم النموذج وغالبًا ما تطبق تصحيحات عالمية قد لا تعالج أخطاء المعايرة المحددة لكل حالة.
لتجاوز هذه التحديات، يقترح المؤلفون نهجًا جديدًا يدمج المعايرة مباشرة في عملية تدريب DNNs، مستخدمين خطأ المعايرة المتوسطة (ACE) كمقياس للمعايرة ودالة خسارة مساعدة قابلة للتفريق. تعالج هذه الطريقة جميع مستويات الثقة بالتساوي، وهو ما يعد مفيدًا بشكل خاص لمهام تقسيم الصور الطبية حيث تكون عدم اليقين أمرًا حاسمًا. يقدمون خسارة قائمة على ACE قابلة للتفريق تستفيد من التنبؤات على مستوى الفوكسل لتوفير تقديرات مستقرة للمعايرة، حتى من الصور الفردية. كما يقدم البحث مخططات موثوقية مجموعة البيانات لتصور أداء المعايرة عبر مجموعات البيانات ويقارن خسارة mL1-ACE المقترحة مع الطرق التقليدية، مما يظهر تحسينات كبيرة في المعايرة مع الحفاظ على أداء التقسيم.
النتائج
في هذا القسم، يقدم المؤلفون نتائج تقييمهم لنماذج مختلفة تم تدريبها على أربعة مجموعات بيانات: ACDC، AMOS، BraTS، وKiTS. حقق النموذج الأساسي، الذي يستخدم معامل تشابه دايس (DSC) وخسائر الانتروبيا المتقاطعة (CE)، درجات DSC تبلغ 0.871، 0.883، 0.905، و0.859، على التوالي، لمجموعات البيانات. تتم مقارنة هذه الدرجات بالقيم الحديثة، مما يشير إلى أنه على الرغم من أن النموذج الأساسي يؤدي بشكل جيد، إلا أن هناك مجالًا للتحسين. أدى إدخال خسائر mL1-ACE المساعدة، مع استراتيجيات الصناديق الصلبة والناعمة، إلى أداء أفضل باستمرار من النموذج الأساسي من حيث ACE المتوسط والميكروي عبر جميع مجموعات البيانات، مع الحفاظ على أداء DSC قابل للمقارنة.
تشير النتائج إلى أن التدريب باستخدام CE فقط يؤدي إلى تحسينات متواضعة في ACE المتوسط، مع تقليل بنسبة حوالي 3-6% عبر مجموعات البيانات، وانخفاض مستمر في DSC، مما يبرز التوازن بين دقة التقسيم والمعايرة الاحتمالية. بالمقابل، تقلل كل من طرق mL1-ACE الصلبة والناعمة بشكل كبير من ACE المتوسط بنسبة 7-22% و16-44%، على التوالي، مع الحفاظ على مستويات DSC مشابهة للنموذج الأساسي DSC+CE. ومن الجدير بالذكر أن mL1-ACE ذات الصناديق الناعمة أظهرت أكبر تحسينات، حيث حققت زيادة بنسبة 46% في ACE المتوسط على ACDC و70% في ACE الميكروي على نفس مجموعة البيانات. تؤكد هذه النتائج فعالية خسائر mL1-ACE المقترحة في تعزيز معايرة النموذج وأداء التقسيم، خاصة مع نهج الصناديق الناعمة.
المناقشة
في قسم المناقشة من الورقة، يستكشف المؤلفون الجوانب الحرجة لتقدير عدم اليقين في تقسيم الصور الطبية، مميزين بين عدم اليقين العشوائي وعدم اليقين المعرفي. يسلطون الضوء على منهجيات مختلفة لنمذجة هذه الشكوك، مثل إسقاط مونت كارلو والفرق العميقة لعدم اليقين المعرفي، وتنبؤات التباين لكل فوكسل لعدم اليقين العشوائي. يؤكد المؤلفون على أهمية التقاط كلا نوعي عدم اليقين لتعزيز المعايرة وموثوقية نماذج التقسيم، خاصة في سياق تحدي QU-BraTS، الذي أظهر أن دمج هذه الطرق يؤدي إلى تحسين الأداء في مهام تقسيم أورام الدماغ. على الرغم من هذه التقدمات، يشير المؤلفون إلى أن العديد من الأعمال الحالية تفتقر إلى نهج صارم لتقدير عدم اليقين والمعايرة، مما يبرز الحاجة إلى مقاييس موحدة.
يتناول المؤلفون أيضًا تقنيات المعايرة، مع تعريف المعايرة على أنها التوافق بين الاحتمالات المتوقعة والدقة الفعلية. يقدمون مخططات موثوقية كأداة لتصور أداء المعايرة ويناقشون مقاييس المعايرة المختلفة، بما في ذلك خطأ المعايرة المتوقع (ECE) وخطأ المعايرة المتوسطة (ACE). يميز القسم أيضًا بين طرق المعايرة بعد الحدث، مثل قياس الحرارة، وتقنيات المعايرة أثناء التدريب، التي تستفيد من العدد الكبير من المعلمات في الشبكات العصبية العميقة (DNNs). يقترح المؤلفون أن دمج خسائر المعايرة المساعدة أثناء التدريب يمكن أن يعزز بشكل كبير معايرة النموذج، خاصة في مهام التقسيم، حيث توفر التنبؤات الكثيفة ثروة من البيانات لتقديرات المعايرة الأكثر استقرارًا. يخلصون إلى أنه على الرغم من أن خسائر المعايرة المساعدة أظهرت وعدًا في تحسين موثوقية النموذج، إلا أن هناك حاجة إلى مزيد من البحث لتحسين هذه التقنيات ومعالجة التحديات الفريدة التي تطرحها المخرجات عالية الأبعاد للتقسيم.
القيود
يتناول قسم القيود القيود المفروضة على استخدام خطأ المعايرة المتوقع (ECE) وخسائر المعايرة في تقييم أداء الشبكات العصبية العميقة (DNNs) في مهام التقسيم الدلالي. يُنتقد ECE لكونه متأثرًا بشكل مفرط بتوزيع العد في الصناديق، خاصة الصناديق الأولى والأخيرة، مما قد يؤدي إلى عدم الحساسية تجاه الأخطاء ذات الصلة سريريًا. تتفاقم هذه المشكلة في المخرجات عالية الأبعاد، مثل التصوير الطبي ثلاثي الأبعاد، حيث تنتمي الغالبية العظمى من الفوكسلات إلى فئة الخلفية. للتخفيف من هذه القيود، يقترح المؤلفون خطأ المعايرة المتوسطة (ACE) كبديل أكثر موثوقية، حيث إنه يوازن بين جميع الصناديق ويؤكد على المعايرة بالقرب من حدود القرار (0.5). على الرغم من أن خطأ المعايرة الأقصى (MCE) يوفر رؤى حول السيناريوهات الأسوأ، إلا أن ACE يُظهر أنه مقياس موثوق للمعايرة بسبب ارتباطه بـ MCE.
بالإضافة إلى ذلك، بينما تحسن خسائر المعايرة المساعدة المستخدمة في الدراسة مقاييس المعايرة، يعترف المؤلفون بأن هذه التحسينات لا تضمن بالضرورة نتائج سريرية أفضل. يعتمد الأسلوب على تصنيف احتمالات المقدمة مما يحد من قدرة النموذج على التقاط الارتباطات المكانية في عدم اليقين والثقة. يمكن أن تستكشف الأعمال المستقبلية تقنيات تصنيف تكيفية لتناسب بشكل أفضل توزيعات الاحتمالات الخاصة بمجموعة البيانات. يؤكد المؤلفون على أهمية تحقيق التوازن بين الدقة والمعايرة، مما يوفر للممارسين السيطرة على أهداف المعايرة المصممة لتطبيقات محددة. يدعم تقييمهم عبر مجموعات بيانات متنوعة موثوقية نهجهم، ويقترحون أن تجارب التعميم عبر مجموعات البيانات يمكن أن تقيم المزيد من القوة في سيناريوهات التوزيع المتغيرة.
DOI: https://doi.org/10.1109/tmi.2026.3673118
PMID: https://pubmed.ncbi.nlm.nih.gov/41811722
Publication Date: 2026-03-11
Author(s): Theodore Barfoot et al.
Primary Topic: Advanced Neural Network Applications
Overview
In this research, the authors address the issue of overconfidence in deep neural networks used for medical image segmentation, which can undermine their reliability in clinical settings. They propose the use of differentiable formulations of marginal L1 Average Calibration Error (mL1-ACE) as an auxiliary loss function, computed on a per-image basis, to enhance pixel-wise calibration. Their experiments across four datasets (ACDC, AMOS, KiTS, BraTS) reveal that incorporating mL1-ACE effectively reduces both Average Calibration Error (ACE) and Maximum Calibration Error (MCE), while maintaining high Dice Similarity Coefficients (DSCs). Notably, the soft-binned variant of mL1-ACE achieves the most significant calibration improvements, albeit with a slight compromise in segmentation performance, while the hard-binned variant preserves segmentation accuracy with moderate calibration enhancements.
The study concludes that integrating calibration-aware losses during training leads to substantial reductions in calibration errors, with the soft-binned mL1-ACE variant reducing macro- and micro-averaged ACE by 33% and 51%, respectively, at a minor cost to segmentation accuracy. In contrast, the hard-binned variant improves calibration by 16% and 32% for macro- and micro-averaged ACE, respectively, while slightly enhancing segmentation accuracy. The authors emphasize the importance of balancing calibration and accuracy, providing practitioners with options tailored to specific clinical needs. They also advocate for the use of Average Calibration Error (ACE) over Expected Calibration Error (ECE) for better assessment of miscalibration, particularly in critical boundary regions. Additionally, they introduce dataset reliability histograms to visualize calibration performance across datasets and highlight challenges related to missing ground-truth classes, suggesting avenues for future research. Overall, their findings underscore the potential for optimizing calibration during training to enhance the reliability of AI systems in medical imaging.
Introduction
In the introduction, the authors address the critical issue of miscalibration in deep neural networks (DNNs), particularly in the context of medical imaging, where overconfidence in predictions can lead to significant risks in clinical decision-making. They emphasize the importance of well-calibrated models, which align predicted probabilities with actual outcomes, thereby enhancing the trustworthiness of AI applications in healthcare. Traditional post-hoc calibration methods, such as Platt scaling and temperature scaling, are noted for their limitations, as they do not fully utilize the model’s learning capacity and often apply global corrections that may not address instance-specific calibration errors.
To overcome these challenges, the authors propose a novel approach that integrates calibration directly into the training process of DNNs, utilizing the Average Calibration Error (ACE) as both a calibration metric and a differentiable auxiliary loss function. This method treats all confidence levels equally, which is particularly beneficial for medical segmentation tasks where uncertainty is critical. They introduce a differentiable ACE-based loss that leverages voxel-level predictions to provide stable calibration estimates, even from single images. The paper also presents dataset reliability histograms for visualizing calibration performance across datasets and compares their proposed mL1-ACE loss with traditional methods, demonstrating significant improvements in calibration while maintaining segmentation performance.
Results
In this section, the authors present the results of their evaluation of various models trained on four datasets: ACDC, AMOS, BraTS, and KiTS. The baseline model, which utilizes Dice Similarity Coefficient (DSC) and cross-entropy (CE) losses, achieved DSC scores of 0.871, 0.883, 0.905, and 0.859, respectively, for the datasets. These scores are compared to state-of-the-art values, indicating that while the baseline performs well, there is room for improvement. The introduction of mL1-ACE auxiliary losses, with both hard and soft binning strategies, consistently outperformed the baseline in terms of macro-averaged ACE and micro-averaged ACE across all datasets, while maintaining comparable DSC performance.
The results indicate that training with CE alone yields only modest improvements in macro-ACE, with a reduction of approximately 3-6% across datasets, and a consistent decrease in DSC, highlighting a trade-off between segmentation accuracy and probabilistic calibration. In contrast, both hard and soft mL1-ACE approaches significantly reduce macro-ACE by 7-22% and 16-44%, respectively, while preserving DSC levels similar to the DSC+CE baseline. Notably, soft-binned mL1-ACE demonstrated the most substantial improvements, achieving a 46% increase in macro-averaged ACE on ACDC and 70% in micro-averaged ACE on the same dataset. These findings underscore the effectiveness of the proposed mL1-ACE losses in enhancing model calibration and segmentation performance, particularly with the soft-binning approach.
Discussion
In the discussion section of the paper, the authors explore the critical aspects of uncertainty estimation in medical image segmentation, distinguishing between aleatoric and epistemic uncertainties. They highlight various methodologies for modeling these uncertainties, such as Monte Carlo dropout and deep ensembles for epistemic uncertainty, and per-voxel variance predictions for aleatoric uncertainty. The authors emphasize the importance of capturing both uncertainty types to enhance the calibration and reliability of segmentation models, particularly in the context of the QU-BraTS challenge, which demonstrated that combining these methods leads to improved performance in brain tumor segmentation tasks. Despite these advancements, the authors note that many existing works lack a rigorous approach to uncertainty estimation and calibration, underscoring the need for standardized metrics.
The authors also delve into calibration techniques, defining calibration as the alignment between predicted probabilities and actual accuracy. They introduce reliability diagrams as a tool for visualizing calibration performance and discuss various calibration metrics, including Expected Calibration Error (ECE) and Average Calibration Error (ACE). The section further distinguishes between post-hoc calibration methods, such as Temperature Scaling, and train-time calibration techniques, which leverage the high parameter count of deep neural networks (DNNs). The authors propose that incorporating auxiliary calibration losses during training can significantly enhance model calibration, particularly in segmentation tasks, where dense predictions provide a wealth of data for more stable calibration estimates. They conclude that while auxiliary calibration losses have shown promise in improving model reliability, further research is needed to refine these techniques and address the unique challenges posed by high-dimensional segmentation outputs.
Limitations
The section on limitations discusses the constraints of using Expected Calibration Error (ECE) and calibration losses in assessing the performance of deep neural networks (DNNs) in semantic segmentation tasks. ECE is criticized for being overly influenced by the distribution of counts in bins, particularly the first and last bins, which can lead to insensitivity towards clinically relevant errors. This issue is exacerbated in high-dimensional outputs, such as 3D medical imaging, where the majority of voxels belong to the background class. To mitigate these limitations, the authors propose Average Calibration Error (ACE) as a more robust alternative, as it equally weights all bins and emphasizes calibration near the decision boundary (0.5). Although Maximum Calibration Error (MCE) offers insights into worst-case scenarios, ACE is shown to be a reliable calibration metric due to its correlation with MCE.
Additionally, while the auxiliary calibration losses employed in the study improve calibration metrics, the authors acknowledge that these improvements do not necessarily guarantee better clinical outcomes. The method’s reliance on binning foreground probabilities limits the model’s ability to capture spatial correlations in uncertainty and confidence. Future work could explore adaptive binning techniques to better accommodate dataset-specific probability distributions. The authors emphasize the importance of balancing accuracy and calibration, providing practitioners with control over calibration objectives tailored to specific applications. Their evaluation across diverse datasets supports the reliability of their approach, and they suggest that cross-dataset generalization experiments could further assess robustness in varying distribution scenarios.
