إطار إحصائي جديد لت quantifying المخاطر والفوائد من أتمتة الذكاء الاصطناعي في تصوير الثدي بالأشعة
A novel statistical framework for quantifying risks and benefits of AI automation in screening mammography

شارك:
المجلة: PLOS Digital Health، المجلد: 5، العدد: 2
DOI: https://doi.org/10.1371/journal.pdig.0001231
PMID: https://pubmed.ncbi.nlm.nih.gov/41746988
تاريخ النشر: 2026-02-26
المؤلف: Zhenyun Du وآخرون
الموضوع الرئيسي: الذكاء الاصطناعي في اكتشاف السرطان

نظرة عامة

تقدم ورقة البحث إطارًا لتحديد العتبات المثلى لاستخدام الذكاء الاصطناعي (AI) كأداة فرز في الأشعة، وخاصة لفحص الماموجرامات. نظرًا للزيادة المتزايدة في حجم التصوير ونقص القوى العاملة، يُقترح استخدام الذكاء الاصطناعي لـ “استبعاد” الحالات منخفضة المخاطر، مما يسمح لأطباء الأشعة بالتركيز على الفحوصات عالية المخاطر. قامت الدراسة بتحليل 114,229 ماموجرام ثنائي الأبعاد رقمي من 2006 إلى 2023، باستخدام نموذج ذكاء اصطناعي مفتوح المصدر، ميراى، لتعيين درجات المخاطر وتقييم عتبات مختلفة لاستبعاد الحالات.

يقوم الإطار بتحديد التوازن بين تقليل عبء العمل والمخاطر التشخيصية المرتبطة بعتبات استبعاد الذكاء الاصطناعي المختلفة. كما يترجم عتبات درجات الذكاء الاصطناعي إلى معدلات خطأ ذات صلة سريريًا ويحسب تخفيضات عبء العمل المقابلة. من المهم أن يميز الإطار بين السرطانات التي قد تُفوت بسبب استبعاد الذكاء الاصطناعي وتلك التي كانت ستفوت تحت الممارسة القياسية. بالإضافة إلى ذلك، يأخذ في الاعتبار التأثير المحتمل على أداء أطباء الأشعة، مشيرًا إلى أن مجموعة قراءة أصغر وأكثر غنى قد تؤدي إلى زيادة اكتشاف السرطان بين الحالات المحتفظ بها. بدلاً من اقتراح عتبة محددة، تقدم الدراسة أداة قابلة للتعميم للممارسات وصانعي السياسات لتقييم استراتيجيات استبعاد الذكاء الاصطناعي بناءً على البيانات المحلية وتحمل المخاطر.

مقدمة

تسلط مقدمة ورقة البحث هذه الضوء على الزيادة المتزايدة في دمج الذكاء الاصطناعي (AI) في الأشعة، خاصة استجابةً للعبء المتزايد الذي يواجهه أطباء الأشعة. تشير البيانات إلى زيادة كبيرة في عبء العمل، حيث شهد أطباء الأشعة في أقسام الطوارئ زيادة بمقدار أربعة أضعاف من 2006 إلى 2020، وزيادة بنسبة 60% في وحدات قيمة العمل المفوترة (RVUs) بين أطباء الأشعة الأكاديميين من 2008 إلى 2020. لقد أدت هذه الزيادة في الطلب، جنبًا إلى جنب مع تدفق ثابت من المقيمين الجدد في الأشعة، إلى تفاقم معدلات الإرهاق وخلق عدم توازن في القوى العاملة. يقدم الذكاء الاصطناعي حلاً محتملاً، خاصة من خلال تطبيقه كأداة فرز أو “استبعاد” يمكن أن تحدد الماموجرامات منخفضة المخاطر، مما يقلل من عدد الحالات التي تتطلب تفسير أطباء الأشعة.

تناقش الورقة وعد الذكاء الاصطناعي في إدارة الماموجرامات بشكل فعال، حيث أن أكثر من 99% من الحالات سلبية للسرطان. من خلال استخدام الذكاء الاصطناعي لاستبعاد هذه الحالات منخفضة المخاطر، يمكن لأطباء الأشعة التركيز على الفحوصات الأكثر اشتباهًا، مما يؤدي إلى توفير كبير في الوقت وتقليل عبء العمل. يؤكد المؤلفون على أهمية إنشاء عتبة فرز مناسبة لخوارزميات الذكاء الاصطناعي، التي تولد درجات مخاطر مستمرة للصور. تهدف الدراسة إلى تطوير إطار كمي لتحديد العتبات المثلى للاستبعاد، وتقييم التوازن بين الفوائد والمخاطر المرتبطة بالعتبات المختلفة. باستخدام درجات مخاطر الذكاء الاصطناعي من نموذج ميراى المطبق على 114,229 ماموجرام، يقترح المؤلفون طرقًا لمحاكاة عتبات الاستبعاد وقياس تأثيرها على تقليل عبء الحالات واكتشاف السرطان، بهدف إبلاغ جهود السياسة والتنفيذ المستقبلية في الأشعة المدعومة بالذكاء الاصطناعي.

الطرق

في هذا القسم، يصف المؤلفون منهجيتهم لتقييم تأثير عتبات “استبعاد” الذكاء الاصطناعي المتغيرة على نتائج الماموجرافيا. استخدموا تقنيات المحاكاة باستخدام درجات المخاطر التي تم إنشاؤها بواسطة نموذج التعلم العميق. يسمح هذا النهج باستكشاف منهجي للتوازنات المعنية في تعديل هذه العتبات، مما يوفر رؤى حول كيفية تأثير الإعدادات المختلفة على الدقة التشخيصية ونتائج المرضى في فحص سرطان الثدي. من المتوقع أن تُعلم النتائج المستخلصة من هذه المحاكاة أفضل الممارسات في تطبيق الذكاء الاصطناعي في البيئات السريرية.

النتائج

يقدم قسم “النتائج” في ورقة البحث النتائج الرئيسية المستمدة من التجارب أو التحليلات التي تم إجراؤها. يوضح نتائج الاختبارات المختلفة، مع تسليط الضوء على العلاقات الإحصائية الهامة والاتجاهات الملحوظة في البيانات. غالبًا ما تكون النتائج مصحوبة بأرقام وجداول ذات صلة توضح النتائج بشكل كمي، مما يسمح بفهم أوضح للتداعيات.

بالإضافة إلى ذلك، قد يناقش القسم أداء النماذج أو الفرضيات المختبرة، موفرًا مقاييس مثل الدقة، والموثوقية، أو معاملات الارتباط. تعتبر هذه النتائج حاسمة للتحقق من أسئلة البحث المطروحة في الدراسة وتساهم في الفهم الأوسع للموضوع قيد التحقيق. بشكل عام، تؤكد النتائج على أهمية البحث وإمكاناته في التطبيقات ذات الصلة في هذا المجال.

المناقشة

في هذا القسم، يناقش المؤلفون الاعتبارات الأخلاقية والتعريفات التشغيلية ذات الصلة بدراستهم حول استخدام الذكاء الاصطناعي في الأشعة، وخاصة في فحص سرطان الثدي. تؤكد الدراسة، التي وافق عليها مجلس مراجعة المؤسسات في UCSF، أن تقييم مقاييس أداء الذكاء الاصطناعي بشكل منفصل غير كافٍ دون النظر في التفاعل بين الذكاء الاصطناعي وأطباء الأشعة. يتم تعريف مقاييس رئيسية مثل معدل تقليل عبء العمل (CRR)، وقيمة التنبؤ السلبية للذكاء الاصطناعي (AI-NPV)، ومعدل الخطأ الإجمالي للذكاء الاصطناعي (G-FOR) لتوضيح التوازن بين استبعاد الحالات وإمكانية التشخيصات المفقودة. على سبيل المثال، يمكن أن يؤدي CRR أعلى، الذي يتم تحقيقه من خلال تعيين عتبة أكثر صرامة للذكاء الاصطناعي، إلى زيادة G-FOR، مما يعني أنه قد يتم استبعاد المزيد من حالات سرطان الثدي بشكل غير صحيح.

يقدم المؤلفون إطار محاكاة باستخدام نموذج الذكاء الاصطناعي ميراى لاستكشاف عتبات “الاستبعاد” المختلفة وتأثيراتها على الدقة التشخيصية وإدارة عبء الحالات. يظهرون أنه بينما يمكن أن تؤدي العتبة الأعلى إلى تقليل عدد الحالات التي يحتاج أطباء الأشعة لمراجعتها، فقد تزيد أيضًا من احتمال فقدان السرطانات. على العكس، قد تؤدي العتبة المنخفضة إلى الاحتفاظ بمزيد من الحالات للمراجعة، مما قد يحسن معدلات الاكتشاف ولكن على حساب زيادة عبء العمل. تهدف الدراسة إلى توفير إطار كمي لتقييم هذه التوازنات، مشددة على أن العتبة المثلى ستعتمد على عوامل مختلفة، بما في ذلك خصائص السكان المحليين والقدرات المؤسسية. تهدف النتائج إلى إبلاغ السياسات والممارسات السريرية بدلاً من أن تكون توصيات نهائية لتنفيذ الذكاء الاصطناعي في الأشعة.

Journal: PLOS Digital Health, Volume: 5, Issue: 2
DOI: https://doi.org/10.1371/journal.pdig.0001231
PMID: https://pubmed.ncbi.nlm.nih.gov/41746988
Publication Date: 2026-02-26
Author(s): Zhenyun Du et al.
Primary Topic: AI in cancer detection

Overview

The research paper presents a framework for determining optimal thresholds for using artificial intelligence (AI) as a triage tool in radiology, specifically for screening mammograms. Given the increasing volume of imaging and workforce shortages, AI is proposed to “rule out” low-risk cases, allowing radiologists to concentrate on higher-risk examinations. The study analyzed 114,229 bilateral 2D digital screening mammograms from 2006 to 2023, utilizing an open-source AI model, Mirai, to assign risk scores and evaluate various thresholds for ruling out cases.

The framework quantifies the trade-offs between workload reduction and diagnostic risk associated with different AI rule-out thresholds. It translates AI score thresholds into clinically relevant error rates and calculates corresponding workload reductions. Importantly, the framework differentiates between cancers that would be missed due to AI rule-out and those that would have been missed under standard practice. Additionally, it considers the potential impact on radiologist performance, suggesting that a smaller, enriched reading pool may lead to increased cancer detection among retained cases. Rather than proposing a specific threshold, the study offers a generalizable tool for practices and policymakers to assess AI rule-out strategies based on local data and risk tolerances.

Introduction

The introduction of this research paper highlights the increasing integration of artificial intelligence (AI) in radiology, particularly in response to the escalating workload faced by radiologists. Data indicates a significant rise in workload, with on-call radiologists in emergency departments experiencing a fourfold increase from 2006 to 2020, and a 60% increase in billed work relative value units (RVUs) among academic radiologists from 2008 to 2020. This surge in demand, coupled with a stagnant influx of new radiology residents, has exacerbated burnout rates and created a workforce imbalance. AI presents a potential solution, especially through its application as a triage or “rule-out” tool that can identify low-risk screening mammograms, thereby reducing the number of cases requiring radiologist interpretation.

The paper discusses the promise of AI in efficiently managing screening mammograms, where over 99% of cases are cancer-negative. By employing AI to rule out these low-risk cases, radiologists can focus on more suspicious exams, leading to significant time savings and workload reduction. The authors emphasize the importance of establishing an appropriate Triage Threshold for AI algorithms, which generate continuous risk scores for images. The study aims to develop a quantitative framework for determining optimal rule-out thresholds, evaluating the trade-offs between benefits and risks associated with different thresholds. Using AI risk scores from the Mirai model applied to 114,229 screening mammograms, the authors propose methods for simulating rule-out thresholds and quantifying their impact on caseload reduction and cancer detection, ultimately aiming to inform future policy and implementation efforts in AI-assisted radiology.

Methods

In this section, the authors describe their methodology for evaluating the impact of varying AI “rule-out” thresholds on screening mammography outcomes. They employed simulation techniques utilizing risk scores generated by a deep learning model. This approach allows for a systematic exploration of the trade-offs involved in adjusting these thresholds, providing insights into how different settings can influence diagnostic accuracy and patient outcomes in breast cancer screening. The findings from these simulations are expected to inform best practices in the application of AI in clinical settings.

Results

The “Results” section of the research paper presents key findings derived from the conducted experiments or analyses. It details the outcomes of various tests, highlighting significant statistical relationships and trends observed in the data. The results are often accompanied by relevant figures and tables that illustrate the findings quantitatively, allowing for a clearer understanding of the implications.

Additionally, the section may discuss the performance of models or hypotheses tested, providing metrics such as accuracy, precision, or correlation coefficients. These findings are critical for validating the research questions posed in the study and contribute to the broader understanding of the topic under investigation. Overall, the results underscore the importance of the research and its potential applications in the relevant field.

Discussion

In this section, the authors discuss the ethical considerations and operational definitions relevant to their study on the use of AI in radiology, specifically in breast cancer screening. The study, approved by the UCSF Institutional Review Board, emphasizes that evaluating AI performance metrics in isolation is inadequate without considering the interaction between AI and radiologists. Key metrics such as the Caseload Reduction Rate (CRR), AI Negative Predictive Value (AI-NPV), and AI Gross False Omission Rate (G-FOR) are defined to illustrate the trade-offs between ruling out cases and the potential for missed diagnoses. For instance, a higher CRR, achieved by setting a more stringent AI threshold, can lead to increased G-FOR, meaning more cases of breast cancer may be incorrectly ruled out.

The authors present a simulation framework using the Mirai AI model to explore various “rule-out” thresholds and their implications on diagnostic accuracy and caseload management. They demonstrate that while a higher threshold can reduce the number of cases radiologists need to review, it may also increase the likelihood of missing cancers. Conversely, a lower threshold may retain more cases for review, potentially improving detection rates but at the cost of increased workload. The study aims to provide a quantitative framework for assessing these trade-offs, emphasizing that the optimal threshold will depend on various factors, including local population characteristics and institutional capacities. The findings are intended to inform policy and clinical practices rather than serve as definitive recommendations for AI implementation in radiology.

شارك: