DOI: https://doi.org/10.1016/j.artmed.2024.102819
PMID: https://pubmed.ncbi.nlm.nih.gov/38553159
تاريخ النشر: 2024-03-01
المؤلف: Federico Cabitza وآخرون
الموضوع الرئيسي: التفاعل بين الإنسان والآلات والسلامة
نظرة عامة
تستكشف هذه الورقة نهجًا جديدًا للذكاء الاصطناعي القابل للتفسير (XAI) يُسمى التفسيرات السابقة، والتي توفر تفسيرات بديلة لكل نتيجة ممكنة بدلاً من تفسير واحد لاحق. هذه الطريقة هي جزء من فئة أوسع تُسمى الذكاء الاصطناعي الاحتكاكي، المصممة لتعزيز فعالية اتخاذ القرار مع التخفيف من المخاطر المعرفية مثل الاعتماد المفرط والتحيز الآلي. تم إجراء دراسة تجريبية شملت 16 جراح عظام لتقييم تأثير التفسيرات السابقة على تشخيص كسور الفقرات في الأشعة السينية. كشفت الدراسة أنه بينما لم يُحسن استخدام التفسيرات السابقة بشكل كبير من دقة التشخيص، إلا أنه زاد بشكل ملحوظ من ثقة الأطباء، وخاصة أولئك الذين لديهم خبرة أقل.
تسلط النتائج الضوء على إمكانية الذكاء الاصطناعي الاحتكاكي لتعزيز تفاعلات أكثر وعيًا مع التكنولوجيا، مما يشجع على اتخاذ قرارات نقدية ومعالجة المخاوف بشأن تدهور المهارات مع مرور الوقت. تتماشى الأبحاث مع مبادئ الذكاء الاصطناعي المتمركز حول الإنسان ومبادئ الصحة الواحدة، مما يبرز أهمية تصورات المستخدمين والتصميم الشامل في أنظمة دعم القرار. على الرغم من أن حجم العينة الصغيرة للدراسة يحد من إمكانية تعميم النتائج، إلا أن أحجام التأثير الكبيرة الملحوظة توفر أساسًا للبحوث المستقبلية، لا سيما في استكشاف كيفية دعم التفسيرات السابقة لمزيد من اتخاذ القرار البشري وتقليل الأخطاء في البيئات السريرية. ستسعى الأعمال المستقبلية إلى تقييم طرق إضافية للاستفادة من مخرجات الذكاء الاصطناعي الأقل استغلالًا على الفور لتعزيز العمليات المعرفية البشرية.
مقدمة
تسلط مقدمة هذه الورقة البحثية الضوء على تطور مبادئ التصميم المتمركز حول المستخدم في أنظمة الكمبيوتر التفاعلية، التي تأثرت بشكل ملحوظ بأعمال نورمان الرائدة، “تصميم الأشياء اليومية.” أكد نورمان على أهمية إنشاء تجارب مستخدمين بديهية وممتعة مع التحذير من جعل الأنظمة بسيطة للغاية، مما قد يؤدي إلى استرخاء المستخدم. تم تطوير هذه الفكرة بشكل أكبر من خلال مفهوم كوبر للاحتكاك المعرفي، الذي يقترح أن بعض مستويات الاحتكاك في اتخاذ القرار يمكن أن تعزز من تفاعل المستخدم وتمنع تدهور المهارات. تفترض الورقة أن أنظمة دعم القرار التي تقدم إجابات حاسمة قد تعزز الاعتماد المفرط بين المستخدمين، مما قد يقلل من مهارات حكمهم.
لمعالجة هذه المخاوف، يقترح المؤلفون نهجًا جديدًا يُسمى “التفسيرات السابقة”، والذي يتضمن تقديم حالات مشابهة للمستخدمين بناءً على أحكامهم الأولية بدلاً من تقديم نصائح مباشرة مولدة بواسطة الآلة. يهدف هذا الأسلوب إلى تشجيع التفكير النقدي والحفاظ على تفاعل المستخدم في عملية اتخاذ القرار. توضح الورقة دراسة تجريبية تركز على تأثير التفسيرات السابقة في البيئات السريرية، وتحديدًا في التشخيص الإشعاعي. تشير النتائج الأولية إلى أنه بينما كانت تحسينات الدقة العامة متواضعة، فإن استخدام الحالات المشابهة أثر بشكل إيجابي على اتخاذ القرار وثقة المستخدم، مما يدل على الفوائد المحتملة لدمج الاحتكاك المعرفي في أنظمة الذكاء الاصطناعي لتعزيز تجربة المستخدم وجودة القرار.
الطرق
في هذه الدراسة، هدف المؤلفون إلى تقييم فعالية تقديم حالات مشابهة للأطباء لتعزيز دقة التشخيص والفائدة المدركة، مع الامتناع عن التصنيف المباشر للحالات الجديدة. شارك ستة عشر طبيبًا، يتكونون من عشرة أطباء جراحة عظام معتمدين من المجلس وستة مقيمين في جراحة العظام، في تقييم 18 حالة أشعة سينية تم اختيارها بناءً على تعقيدها، بما في ذلك أمثلة إيجابية (مع كسور فقرية) وسلبية (بدون كسور). قام كل طبيب بتقييم حالة واحدة في كل مرة عبر استبيان عبر الإنترنت، مقدمًا آراء تشخيصية أولية وتقييمات لصعوبة الحالة وثقته على مقياس ترتيبي من ست نقاط.
شملت المنهجية نظام ذكاء اصطناعي استرجع حالتين الأكثر تشابهًا بناءً على تشابه كوساين، الذي أظهر سابقًا ارتباطًا قويًا مع تقييمات البشر. بعد مراجعة هذه الحالات المشابهة، أشار الأطباء إلى تشخيصاتهم النهائية ومستويات ثقتهم. قاست الدراسة دقة الأساس (قبل الدعم) ودقة ما بعد الدعم، مع تعريف دعم الذكاء الاصطناعي بأنه استرجاع حالات مشابهة دون تقديم نصائح تصنيفية. شملت التحليلات 288 تشخيصًا، مع تطبيق اختبار فرضيات غير معلمي عند مستوى ثقة 95%، وحساب عدد القرارات المطلوبة (NDN) لتحقيق تأثير كبير على اتخاذ القرار. يحدد مقياس NDN، المستمد من مبادئ وبائية، العدد المتوسط من القرارات المطلوبة مع دعم الذكاء الاصطناعي لضمان قرار صحيح واحد، مما يبرز التأثير المحتمل للتفسيرات السابقة على دقة التشخيص.
المناقشة
في هذا القسم، يقدم المؤلفون مفهوم “الذكاء الاصطناعي الاحتكاكي”، الذي يهدف إلى دمج الاحتكاك المعرفي في أنظمة الذكاء الاصطناعي لتعزيز تفاعل المستخدمين بشكل أكثر وعيًا ونقديًا. يتناقض هذا النهج مع التصاميم التقليدية التي تعطي الأولوية للتفاعلات السلسة، مما يؤدي غالبًا إلى الاعتماد المفرط على الأنظمة الآلية. مستندين إلى نظرية النظامين لكاهنمان، يجادل المؤلفون بأن أنظمة الذكاء الاصطناعي يمكن تصميمها لتشجيع التفكير من النظام 2—التحليلي والتأملي—من خلال تضمين احتكاكات تصميمية متعمدة، مثل الحدود الدقيقة التي تبطئ من عمليات اتخاذ القرار. يتم وضع هذا في مواجهة “الأنماط المظلمة”، التي تستغل التحيزات المعرفية للتلاعب بسلوك المستخدم.
يقدم المؤلفون تصنيفًا لتطبيقات الذكاء الاصطناعي الاحتكاكي، بما في ذلك البروتوكولات الحذرة التي تقدم خيارات متعددة أو تمتنع عن تقديم توصيات، والبروتوكولات القضائية التي تقدم حججًا متعارضة لتحدي ثقة المستخدم. يؤكدون على أن مثل هذه التصاميم يمكن أن تخفف من التحيزات وتعزز اتخاذ القرار من خلال تعزيز التأمل بدلاً من الاسترخاء. تظهر الدراسات التجريبية التي أجريت في سياقات اتخاذ القرار الطبي أن الذكاء الاصطناعي الاحتكاكي، وخاصة من خلال التفسيرات السابقة، يمكن أن يحسن دقة التشخيص وثقة المستخدم، خاصة بين الأطباء الأقل خبرة. تشير النتائج إلى أنه بينما يُنظر إلى الاحتكاك غالبًا بشكل سلبي في تفاعل الإنسان مع الكمبيوتر، يمكن استخدامه استراتيجيًا لتعزيز عمليات اتخاذ القرار وتقليل المخاطر المرتبطة بالاعتماد المفرط على التكنولوجيا. يُشجع على إجراء بحوث مستقبلية لاستكشاف تطبيقات إضافية للذكاء الاصطناعي الاحتكاكي لدعم المستخدمين البشريين في سياقات اتخاذ القرار المختلفة.
DOI: https://doi.org/10.1016/j.artmed.2024.102819
PMID: https://pubmed.ncbi.nlm.nih.gov/38553159
Publication Date: 2024-03-01
Author(s): Federico Cabitza et al.
Primary Topic: Human-Automation Interaction and Safety
Overview
This paper investigates a novel approach to explainable AI (XAI) termed pro-hoc explanations, which provide alternative explanations for each possible outcome rather than a single post-hoc explanation. This method is part of a broader category called frictional AI, designed to enhance decision-making effectiveness while mitigating cognitive risks such as over-reliance and automation bias. An empirical study involving 16 orthopedists was conducted to assess the impact of pro-hoc explanations on the diagnosis of vertebral fractures in x-rays. The study revealed that while the use of pro-hoc explanations did not significantly improve diagnostic accuracy, it notably increased the confidence of physicians, particularly those with less experience.
The findings highlight the potential of frictional AI to foster more mindful interactions with technology, encouraging critical decision-making and addressing concerns about skill degradation over time. The research aligns with Human-Centered Artificial Intelligence and One Health principles, emphasizing the importance of user perceptions and holistic design in decision support systems. Although the study’s small sample size limits generalizability, the substantial effect sizes observed provide a foundation for future research, particularly in exploring how pro-hoc explanations can further support human decision-making and reduce errors in clinical settings. Future work will aim to evaluate additional methods for leveraging less immediately exploitable AI outputs to enhance human cognitive processes.
Introduction
The introduction of this research paper highlights the evolution of user-centered design principles in interactive computer systems, notably influenced by Norman’s seminal work, “The Design of Everyday Things.” Norman emphasized the importance of creating intuitive and enjoyable user experiences while cautioning against making systems overly simplistic, which could lead to user complacency. This notion was further developed by Cooper’s concept of cognitive friction, which suggests that some level of friction in decision-making can enhance user engagement and prevent deskilling. The paper posits that decision support systems providing definitive answers may foster over-dependence among users, potentially diminishing their judgment skills.
To address these concerns, the authors propose a novel approach termed “pro-hoc explanations,” which involves presenting users with similar cases based on their initial judgments rather than providing direct machine-generated advice. This method aims to encourage critical thinking and maintain user engagement in the decision-making process. The paper outlines an empirical study focusing on the impact of pro-hoc explanations in clinical settings, specifically in radiological diagnosis. Preliminary findings suggest that while the overall accuracy improvement was modest, the use of similar cases positively influenced decision-making and user confidence, indicating the potential benefits of integrating cognitive friction into AI systems to enhance user experience and decision quality.
Methods
In this study, the authors aimed to evaluate the effectiveness of providing physicians with similar cases to enhance diagnostic accuracy and perceived utility, while refraining from direct classification of new cases. Sixteen physicians, comprising ten board-certified orthopaedic spine subspecialists and six orthopaedic residents, participated in annotating 18 x-ray cases selected for their complexity, including both positive (with vertebral fractures) and negative (without fractures) examples. Each physician assessed one case at a time via an online questionnaire, providing initial diagnostic opinions and ratings of case difficulty and confidence on a six-point ordinal scale.
The methodology involved an AI system that retrieved the two most similar cases based on Cosine similarity, which had previously shown a strong correlation with human ratings. After reviewing these similar cases, physicians indicated their final diagnoses and confidence levels. The study measured baseline accuracy (pre-support) and post-support accuracy, with AI support being defined as the retrieval of similar cases without offering categorical advice. The analysis included 288 diagnoses, applying non-parametric hypothesis testing at a 95% confidence level, and calculated the Number of Decisions Needed (NDN) to achieve a significant effect on decision-making. The NDN metric, derived from epidemiological principles, quantifies the average number of decisions required with AI support to ensure one correct decision, highlighting the potential impact of pro-hoc explanations on diagnostic accuracy.
Discussion
In this section, the authors introduce the concept of “frictional AI,” which aims to incorporate cognitive friction into AI systems to promote more mindful and critical user engagement. This approach contrasts with traditional designs that prioritize seamless interactions, often leading to over-reliance on automated systems. Drawing on Kahneman’s dual-system theory, the authors argue that AI systems can be designed to encourage System 2 thinking—analytical and reflective—by embedding intentional design frictions, such as micro-boundaries that slow down decision-making processes. This is juxtaposed against “dark patterns,” which exploit cognitive biases to manipulate user behavior.
The authors present a typology of frictional AI applications, including cautious protocols that provide multiple options or abstain from giving recommendations, and judicial protocols that present conflicting arguments to challenge user confidence. They emphasize that such designs can mitigate biases and enhance decision-making by fostering reflection rather than complacency. Empirical studies conducted in medical decision-making contexts demonstrate that frictional AI, particularly through pro-hoc explanations, can improve diagnostic accuracy and user confidence, especially among less experienced clinicians. The findings suggest that while friction is often viewed negatively in human-computer interaction, it can be strategically employed to enhance decision-making processes and reduce the risks associated with over-reliance on technology. Future research is encouraged to explore further applications of frictional AI to support human users in various decision-making contexts.
