لضمان اعتماد سريري، تتطلب نماذج الذكاء الاصطناعي تقييم تنفيذ متعدد الأبعاد
To warrant clinical adoption AI models require a multi-faceted implementation evaluation

المجلة: npj Digital Medicine، المجلد: 7، العدد: 1
DOI: https://doi.org/10.1038/s41746-024-01064-1
PMID: https://pubmed.ncbi.nlm.nih.gov/38448743
تاريخ النشر: 2024-03-06
المؤلف: Davy van de Sande وآخرون
الموضوع الرئيسي: أنظمة الصحة، التقييمات الاقتصادية، جودة الحياة

نظرة عامة

تسلط الورقة الضوء على الفجوة بين التقدم السريع في تكنولوجيا الذكاء الاصطناعي (AI) واعتمادها السريري المحدود وترجمة قيمتها في بيئات الرعاية الصحية. أظهر استعراض نتائج التنفيذ في التجارب السريرية العشوائية التي تقيم دعم القرار السريري القائم على الذكاء الاصطناعي تكاملًا غير كافٍ لهذه النتائج، والتي تعتبر حاسمة لفهم العوامل التي تؤثر على نجاح أو فشل أنظمة الذكاء الاصطناعي. يجادل المؤلفون بضرورة اتباع نهج منهجي لدمج نتائج التنفيذ طوال دورة حياة الذكاء الاصطناعي السريرية، بدءًا من التطوير قبل السريري إلى التطبيق بجانب السرير.

تؤكد الخاتمة على ضرورة سد الفجوة بين تطوير نماذج الذكاء الاصطناعي وإقامة ذكاء اصطناعي موثوق ومسؤول في الممارسة السريرية. من خلال إعطاء الأولوية لدمج علم التنفيذ ضمن مبادرات الذكاء الاصطناعي، يؤكد المؤلفون أنه من الممكن تحويل توقعات الذكاء الاصطناعي إلى تدخلات بشرية فعالة، مما يعزز الثقة في تقنيات الذكاء الاصطناعي. يُعتبر هذا النهج ضروريًا لتحقيق الإمكانات التحويلية للذكاء الاصطناعي في تقديم الرعاية الصحية وتعزيز نتائج المرضى.

نقاش

يسلط النقاش الضوء على التحديات والقيود الحالية في تقييم تطبيقات الذكاء الاصطناعي (AI) في الرعاية الصحية، لا سيما فيما يتعلق بأنظمة دعم القرار السريري القائمة على الذكاء الاصطناعي (AICDSS). على الرغم من الأداء الإحصائي الواعد للعديد من AICDSS، فإن أقل من 2% تتقدم إلى ما بعد مرحلة النماذج الأولية، ونحو نصف التجارب السريرية العشوائية (RCTs) التي تم تحليلها لم تظهر تحسينات كبيرة في نتائج المرضى. أظهر تحليل 104 RCTs أنه بينما كانت “الولاء” هي النتيجة الأكثر تكرارًا المبلغ عنها في التنفيذ، غالبًا ما تم تجاهل جوانب حاسمة أخرى مثل “القبول” و”الملاءمة”. يشير هذا إلى وجود فجوة في فهم الآثار الأوسع لتنفيذ AICDSS، لا سيما في البيئات المعقدة للمستشفيات.

يدعو المؤلفون إلى اتباع نهج تقييم متعدد الأبعاد يتضمن كل من مقاييس الفعالية ونتائج التنفيذ، كما هو موضح في الأطر المعتمدة مثل الإطار الموحد لأبحاث التنفيذ (CFIR). يؤكدون على الحاجة إلى الإبلاغ المنهجي عن نتائج التنفيذ لتعزيز الثقة والشفافية في تطبيقات الذكاء الاصطناعي. علاوة على ذلك، يقترحون اعتماد تصاميم هجينة للفعالية والتنفيذ في التجارب السريرية العشوائية المستقبلية لالتقاط ديناميكيات دمج الذكاء الاصطناعي بشكل أفضل في البيئات السريرية. في النهاية، يُعتبر سد الفجوة بين تطوير نماذج الذكاء الاصطناعي وتطبيقها العملي في الرعاية الصحية أمرًا ضروريًا لتحقيق الإمكانات الكاملة للذكاء الاصطناعي لتعزيز رعاية المرضى ونتائجهم.

Journal: npj Digital Medicine, Volume: 7, Issue: 1
DOI: https://doi.org/10.1038/s41746-024-01064-1
PMID: https://pubmed.ncbi.nlm.nih.gov/38448743
Publication Date: 2024-03-06
Author(s): Davy van de Sande et al.
Primary Topic: Health Systems, Economic Evaluations, Quality of Life

Overview

The paper highlights the disparity between rapid advancements in artificial intelligence (AI) technology and its limited clinical adoption and value translation in healthcare settings. A review of implementation outcomes in randomized controlled trials assessing AI-based clinical decision support revealed insufficient integration of these outcomes, which are crucial for understanding the factors influencing the success or failure of AI systems. The authors argue for a systematic approach to incorporate implementation outcomes throughout the clinical AI lifecycle, starting from preclinical development to bedside application.

The conclusion emphasizes the necessity of bridging the gap between AI model development and the establishment of trustworthy, responsible AI in clinical practice. By prioritizing the integration of implementation science within AI initiatives, the authors assert that it is possible to convert AI predictions into effective human interventions, thereby fostering trust in AI technologies. This approach is deemed essential for realizing the transformative potential of AI in healthcare delivery and enhancing patient outcomes.

Discussion

The discussion highlights the current challenges and limitations in the evaluation of artificial intelligence (AI) applications in healthcare, particularly regarding AI-based clinical decision support systems (AICDSS). Despite the promising statistical performance of many AICDSS, less than 2% progress beyond the prototyping stage, and nearly half of the randomized controlled trials (RCTs) analyzed did not demonstrate significant improvements in patient outcomes. The analysis of 104 RCTs revealed that while ‘fidelity’ was the most frequently reported implementation outcome, other critical aspects such as ‘acceptability’ and ‘appropriateness’ were often overlooked. This suggests a gap in understanding the broader implications of AICDSS implementation, particularly in complex hospital environments.

The authors advocate for a multi-faceted evaluation approach that incorporates both effectiveness metrics and implementation outcomes, as outlined in established frameworks like the Consolidated Framework for Implementation Research (CFIR). They emphasize the need for systematic reporting of implementation outcomes to foster trust and transparency in AI applications. Furthermore, they propose the adoption of effectiveness-implementation hybrid designs in future RCTs to better capture the dynamics of AI integration in clinical settings. Ultimately, bridging the gap between AI model development and its practical application in healthcare is essential for realizing the full potential of AI to enhance patient care and outcomes.