DOI: https://doi.org/10.1038/s41598-024-82501-9
PMID: https://pubmed.ncbi.nlm.nih.gov/39730794
تاريخ النشر: 2024-12-28
المؤلف: Julian Senoner وآخرون
الموضوع الرئيسي: الذكاء الاصطناعي القابل للتفسير (XAI)
نظرة عامة
تبحث هذه الدراسة في تأثير الذكاء الاصطناعي القابل للتفسير (AI) على التعاون بين الإنسان والذكاء الاصطناعي، خاصة في المهام التي تتطلب الفحص البصري. تتناول الدراسة التحدي الذي تطرحه أنظمة الذكاء الاصطناعي غير الشفافة، التي تخفي الأسباب وراء توقعاتها، مما يجعل من الصعب على المستخدمين التحقق من هذه التوقعات مقابل خبراتهم. يفترض المؤلفون أن دمج الذكاء الاصطناعي القابل للتفسير، تحديدًا من خلال خرائط الحرارة البصرية، يعزز أداء المهام من خلال تمكين المستخدمين من فهم نتائج الذكاء الاصطناعي بشكل أفضل وسياقها.
لاختبار هذه الفرضية، تم إجراء تجربتين مسجلتين مسبقًا تضمنت خبراء في المجال: عمال المصانع الذين يقيمون المنتجات الإلكترونية بحثًا عن عيوب وأطباء الأشعة الذين يقيمون صور الأشعة السينية للصدر بحثًا عن آفات الرئة. تشير النتائج إلى أن المشاركين الذين استخدموا الذكاء الاصطناعي القابل للتفسير مع خرائط الحرارة تفوقوا على أولئك الذين اعتمدوا على الذكاء الاصطناعي غير الشفاف، محققين تحسينات قدرها 7.7 نقطة مئوية (95% CI: 3.3% إلى 12.0%، P = 0.001) في التصنيع و4.7 نقطة مئوية (95% CI: 1.1% إلى 8.3%، P = 0.010) في المهام الطبية. تؤكد هذه النتائج فعالية الذكاء الاصطناعي القابل للتفسير في تعزيز دقة اتخاذ القرار وتبرز إمكانيته في تحسين موثوقية التعاون بين الإنسان والذكاء الاصطناعي عبر مجالات مختلفة.
الطرق
في هذا القسم، يوضح المؤلفون المنهجية المستخدمة في بحثهم، والتي تضمنت تجربتين عشوائيتين أجريتا في سياقات فحص بصري متميزة: التصنيع (الدراسة 1) والطب (الدراسة 2). تم تكليف المشاركين في كلا الدراستين بأداء أنشطة الفحص البصري. تم تسجيل الفرضيات لكل دراسة مسبقًا، مع توفير روابط للشفافية (الدراسة 1: https://osf.io/7djxb؛ الدراسة 2: https://osf.io/69yqt).
حصل تصميم البحث على موافقة أخلاقية من لجنة الأخلاقيات في ETH زيورخ (EK 2021-N-34)، مما يضمن أن جميع الإجراءات تتماشى مع الإرشادات واللوائح الأخلاقية ذات الصلة. علاوة على ذلك، تم الحصول على موافقة مستنيرة من جميع المشاركين، مما يؤكد النزاهة الأخلاقية للدراسة.
النتائج
في هذا القسم، يقدم المؤلفون نتائج تجربتين عشوائيتين تهدفان إلى تقييم تأثير الذكاء الاصطناعي القابل للتفسير على أداء المهام في التعاون بين الإنسان والذكاء الاصطناعي ضمن سياقات التصنيع والطب. في الدراسة 1، قام عمال المصانع بأداء فحوصات بصرية للمنتجات الإلكترونية لتحديد العيوب، بينما في الدراسة 2، قام أطباء الأشعة بتقييم صور الأشعة السينية للصدر بحثًا عن آفات الرئة المرئية. تم تقسيم المشاركين إلى مجموعتين: واحدة تستخدم ذكاءً اصطناعيًا غير شفاف يوفر درجة غير واضحة، والأخرى تستخدم ذكاءً اصطناعيًا قابلاً للتفسير، والذي قدم نفس الدرجة مع تفسير خرائط الحرارة.
خدمة خرائط الحرارة، على الرغم من عدم توفير معلومات إضافية تتجاوز درجة الذكاء الاصطناعي، ساهمت في تعزيز فهم المستخدم من خلال تسليط الضوء بصريًا على مناطق الاهتمام، مثل العيوب أو الآفات. افترض المؤلفون أن تضمين الذكاء الاصطناعي القابل للتفسير سيؤدي إلى تحسين أداء المهام بين خبراء المجال المشاركين في هذه المهام التعاونية. من المتوقع أن تساهم النتائج من هذه التجارب في فهم كيفية تسهيل الذكاء الاصطناعي القابل للتفسير اتخاذ قرارات أفضل في البيئات المهنية.
المناقشة
في هذا البحث، أجرى المؤلفون تجربتين لتقييم تأثير الذكاء الاصطناعي القابل للتفسير (XAI) على أداء المهام في مهام الفحص البصري عبر مجالات التصنيع والطب. في الدراسة الأولى، قام عمال المصانع بتقييم المنتجات الإلكترونية بحثًا عن العيوب باستخدام إما نظام ذكاء اصطناعي غير شفاف أو نظام ذكاء اصطناعي قابل للتفسير يوفر خرائط حرارة بصرية. أشارت النتائج إلى أن العمال الذين استخدموا الذكاء الاصطناعي القابل للتفسير حققوا دقة متوازنة قدرها 96.3%، أعلى بكثير من دقة 88.6% لأولئك الذين استخدموا الذكاء الاصطناعي غير الشفاف. تم عزو هذا التحسن إلى الالتزام الأفضل بتوقعات الذكاء الاصطناعي الدقيقة وتقليل ملحوظ في التجاوزات الخاطئة للتقييمات الصحيحة للذكاء الاصطناعي. بالمثل، في الدراسة الثانية التي شملت أطباء الأشعة الذين يفحصون صور الأشعة السينية للصدر، تفوق أولئك المدعومون بالذكاء الاصطناعي القابل للتفسير على نظرائهم الذين استخدموا الذكاء الاصطناعي غير الشفاف، محققين دقة متوازنة قدرها 83.8% مقارنة بـ 79.1%.
تؤكد النتائج على أهمية قابلية تفسير الذكاء الاصطناعي في تعزيز اتخاذ القرار البشري. لم يحسن استخدام خرائط الحرارة أداء المهام فحسب، بل فعل ذلك أيضًا دون زيادة الوقت المستغرق في التقييمات، مما يشير إلى أن الذكاء الاصطناعي القابل للتفسير يمكن أن يعزز بشكل فعال قدرات الإنسان في المهام الحرجة. يجادل المؤلفون بأن نتائجهم توفر أدلة قوية على فوائد الذكاء الاصطناعي القابل للتفسير في التطبيقات الواقعية، مما يبرز الحاجة إلى تصميم أنظمة ذكاء اصطناعي تسهل التعاون الأفضل بين الإنسان والذكاء الاصطناعي.
القيود
تقدم الدراسة عدة قيود قد تؤثر على إمكانية تعميم نتائجها. أولاً، ركزت الدراسة بشكل حصري على إعداد عمل إنساني-ذكاء اصطناعي محدد – مهام الفحص البصري – باستخدام شكل معين من القابلية للتفسير من خلال خرائط الحرارة للإشارة إلى العيوب المحتملة في الجودة أو آفات الرئة. بينما تُستخدم خرائط الحرارة بشكل شائع في تفسير صور الذكاء الاصطناعي، قد لا تكون النتائج قابلة للتطبيق على طرق الذكاء الاصطناعي القابل للتفسير الأخرى، والتي قد تؤدي إلى تأثيرات مختلفة على أداء المهام. بالإضافة إلى ذلك، تعترف الدراسة بالضعف الفطري للذكاء الاصطناعي القابل للتفسير، مثل القابلية للهجمات العدائية وإمكانية توليد تفسيرات غير موثوقة قد تضلل المستخدمين.
حد آخر ملحوظ هو انتشار العيوب في الجودة أو آفات الرئة في الصور التجريبية، والذي كان 14%، وهو أعلى بكثير مما يواجهه خبراء المجال عادة. قد يكون لهذا التباين تأثير على توقعات المشاركين أثناء المهمة. ومع ذلك، نظرًا لأن كلا مجموعتي العلاج تعرضتا لنفس مجموعة الصور، يقترح الباحثون أن هذا العامل كان له تأثير ضئيل على النتائج العامة. على الرغم من هذه القيود، يبقى المؤلفون متفائلين بأن تحسينات الأداء من الذكاء الاصطناعي القابل للتفسير يمكن أن تتحقق في سياقات متنوعة حيث تعمل التفسيرات كأدوات مساعدة في اتخاذ القرار.
DOI: https://doi.org/10.1038/s41598-024-82501-9
PMID: https://pubmed.ncbi.nlm.nih.gov/39730794
Publication Date: 2024-12-28
Author(s): Julian Senoner et al.
Primary Topic: Explainable Artificial Intelligence (XAI)
Overview
This research investigates the impact of explainable artificial intelligence (AI) on human-AI collaboration, particularly in tasks requiring visual inspection. The study addresses the challenge posed by black-box AI systems, which obscure the rationale behind their predictions, making it difficult for users to validate these predictions against their expertise. The authors hypothesize that integrating explainable AI, specifically through visual heatmaps, enhances task performance by enabling users to better understand and contextualize AI outputs.
To test this hypothesis, two preregistered experiments were conducted involving domain experts: factory workers assessing electronic products for defects and radiologists evaluating chest X-ray images for lung lesions. The results indicate that participants using explainable AI with heatmaps outperformed those relying on black-box AI, achieving improvements of 7.7 percentage points (95% CI: 3.3% to 12.0%, P = 0.001) in manufacturing and 4.7 percentage points (95% CI: 1.1% to 8.3%, P = 0.010) in medical tasks. These findings underscore the effectiveness of explainable AI in enhancing decision-making accuracy and highlight its potential to improve the reliability of human-AI collaboration across various domains.
Methods
In this section, the authors outline the methodology employed in their research, which involved two randomized experiments conducted in distinct visual inspection contexts: manufacturing (Study 1) and medicine (Study 2). Participants in both studies were tasked with performing visual inspection activities. The hypotheses for each study were preregistered, with links provided for transparency (Study 1: https://osf.io/7djxb; Study 2: https://osf.io/69yqt).
The research design received ethical approval from the Ethics Commission of ETH Zurich (EK 2021-N-34), ensuring that all procedures adhered to relevant ethical guidelines and regulations. Furthermore, informed consent was obtained from all participants, affirming the ethical integrity of the study.
Results
In this section, the authors present the results of two randomized experiments aimed at evaluating the impact of explainable AI on task performance in human-AI collaboration within manufacturing and medical contexts. In Study 1, factory workers performed visual inspections of electronic products to identify defects, while in Study 2, radiologists assessed chest X-ray images for visible lung lesions. Participants were divided into two groups: one utilizing a black-box AI that provided an opaque score, and the other using explainable AI, which offered the same score along with a heatmap explanation.
The heatmap, while not providing additional information beyond the AI score, served to enhance user understanding by visually highlighting areas of interest, such as defects or lesions. The authors hypothesized that the inclusion of explainable AI would lead to improved task performance among domain experts engaged in these collaborative tasks. The findings from these experiments are expected to contribute to the understanding of how explainable AI can facilitate better decision-making in professional settings.
Discussion
In this research, the authors conducted two experiments to evaluate the impact of explainable AI (XAI) on task performance in visual inspection tasks across manufacturing and medical domains. In the first study, factory workers assessed electronic products for defects using either a black-box AI system or an explainable AI system that provided visual heatmaps. Results indicated that workers using explainable AI achieved a balanced accuracy of 96.3%, significantly higher than the 88.6% accuracy of those using black-box AI. This improvement was attributed to better adherence to accurate AI predictions and a notable reduction in erroneous overrides of correct AI assessments. Similarly, in the second study involving radiologists inspecting chest X-ray images, those supported by explainable AI outperformed their peers with black-box AI, achieving a balanced accuracy of 83.8% compared to 79.1%.
The findings underscore the importance of AI explainability in enhancing human decision-making. The use of heatmaps not only improved task performance but also did so without increasing the time taken for assessments, suggesting that explainable AI can effectively augment human capabilities in critical tasks. The authors argue that their results provide robust evidence for the benefits of explainable AI in real-world applications, emphasizing the need for designing AI systems that facilitate better human-AI collaboration.
Limitations
The research presents several limitations that may affect the generalizability of its findings. Firstly, the study focused exclusively on a specific human-AI work setting—visual inspection tasks—utilizing a particular form of explainability through heatmaps to indicate potential quality defects or lung lesions. While heatmaps are commonly employed in AI image interpretation, the results may not be applicable to other explainable AI methods, which could yield different impacts on task performance. Additionally, the study acknowledges the inherent vulnerabilities of explainable AI, such as susceptibility to adversarial attacks and the potential for generating unreliable explanations that could mislead users.
Another notable limitation is the prevalence of quality defects or lung lesions in the experimental images, which was 14%, significantly higher than what domain experts typically encounter. This discrepancy may have influenced participants’ expectations during the task. However, since both treatment groups were exposed to the same set of images, the researchers suggest that this factor likely had a minimal effect on the overall findings. Despite these limitations, the authors remain optimistic that performance enhancements from explainable AI could be realized in various contexts where explanations function as decision aids.
