العيوب الخفية وراء دقة المستوى الخبير لرؤية GPT-4 متعددة الوسائط في الطب
Hidden flaws behind expert-level accuracy of multimodal GPT-4 vision in medicine

شارك:
المجلة: npj Digital Medicine، المجلد: 7، العدد: 1
DOI: https://doi.org/10.1038/s41746-024-01185-7
PMID: https://pubmed.ncbi.nlm.nih.gov/39043988
تاريخ النشر: 2024-07-23
المؤلف: Qiao Jin وآخرون
الموضوع الرئيسي: الذكاء الاصطناعي في الرعاية الصحية والتعليم

نظرة عامة

تدرس الدراسة أداء نموذج المحول المدرب مسبقًا التوليدي 4 مع الرؤية (GPT-4V) في مهام التشخيص الطبي، من خلال تطبيقه في تحديات الصور في مجلة نيو إنجلاند الطبية (NEJM). بينما أبرزت الأبحاث السابقة تفوق GPT-4V على الأطباء البشر في دقة الأسئلة متعددة الخيارات، توسع هذه الدراسة التقييم ليشمل مبررات النموذج لفهم الصور، واسترجاع المعرفة الطبية، والتفكير متعدد الوسائط. تشير النتائج إلى أن GPT-4V يحقق دقة متعددة الخيارات تبلغ 81.6%، متجاوزًا دقة الأطباء البشر البالغة 77.8%، ويظهر دقة تزيد عن 78% في السيناريوهات التي يقدم فيها الأطباء إجابات غير صحيحة.

على الرغم من هذه النتائج الواعدة، تكشف الدراسة عن مخاوف كبيرة بشأن جودة مبررات GPT-4V، لا سيما في فهم الصور، حيث لوحظت أخطاء في التفكير في 35.5% من الإجابات النهائية الصحيحة، مع نسبة 27.2% تُعزى تحديدًا إلى أخطاء في فهم الصور. تؤكد هذه النتائج على أهمية إجراء تقييمات أكثر شمولاً للمبررات التي ينتجها GPT-4V قبل دمجه المحتمل في البيئات السريرية، مما يبرز الحاجة إلى الحذر في الاعتماد فقط على دقته العالية في التقييمات متعددة الخيارات.

الطرق

يستعرض قسم “الطرق” التصميم التجريبي والتقنيات التحليلية المستخدمة في الدراسة. استخدم الباحثون نهجًا كميًا، حيث تم استخدام التحليلات الإحصائية لتقييم البيانات التي تم جمعها من تجارب مختلفة. تضمنت المنهجيات المحددة تطبيق نماذج الانحدار لتقييم العلاقة بين المتغيرات المستقلة والتابعة، مع ضمان نتائج قوية من خلال الضوابط المناسبة والتوزيع العشوائي.

بالإضافة إلى ذلك، دمجت الدراسة تقنيات حسابية متقدمة لتحليل مجموعات بيانات كبيرة، مما سمح بتحديد الأنماط والارتباطات المهمة. ساعد استخدام أدوات البرمجيات لتصور البيانات في تسهيل تفسير النتائج، مما عزز وضوح النتائج. بشكل عام، أسس الإطار المنهجي أساسًا قويًا لاستنتاجات الدراسة، مما يضمن موثوقية وصلاحية النتائج المبلغ عنها.

المناقشة

في هذا القسم، يناقش المؤلفون المنهجية والنتائج من تقييمهم لتحدي صور NEJM، الذي شمل 207 سؤالًا حديثًا في التصوير الطبي. تم تصنيف هذه الأسئلة إلى ثلاث مستويات من الصعوبة بناءً على أداء المستخدم: “سهل” (56-88% صحيح)، “متوسط” (44-56%)، و”صعب” (28-44%). استخدم المؤلفون GPT-4V وBiomedCLIP لتقييم التحديات، حيث تم تحفيز GPT-4V لتحليل الصور وتقديم مبررات لإجاباته. تم اختبار BiomedCLIP، وهو نموذج لغوي متعدد الوسائط، في إعداد بدون تدريب مسبق، محققًا دقة منخفضة تبلغ 25.1%، مما يدل على التحدي الكبير الذي تطرحه هذه الأسئلة على نماذج الرؤية-اللغة الأصغر.

شمل عملية التقييم تصنيفًا من مرحلتين من قبل محترفين طبيين، الذين قاموا بتقييم استجابات GPT-4V عبر ثلاث فئات: فهم الصورة، واسترجاع المعرفة الطبية، والتفكير خطوة بخطوة. تم تصنيف كل استجابة على أنها “صحيحة”، “جزئيًا صحيحة”، أو “غير صحيحة”، مع توفير إرشادات مفصلة لكل فئة. كانت اتفاقية المقيمين بين المصنفين متوسطة إلى كبيرة، حيث تراوحت قيم كابا كوهين من 0.470 إلى 0.698. من الجدير بالذكر أن GPT-4V تفوق أو تساوى مع الأطباء البشر في ستة من تسع تخصصات، وكان هناك ارتباط قوي بين القدرة على تحديد العيوب في مبررات GPT-4V ودقة إجابات الأطباء. تجمع هذه الدراسة بشكل فريد بين مقاييس الأداء الكمية والتقييمات المنهجية لعمليات التفكير وراء مخرجات النموذج.

Journal: npj Digital Medicine, Volume: 7, Issue: 1
DOI: https://doi.org/10.1038/s41746-024-01185-7
PMID: https://pubmed.ncbi.nlm.nih.gov/39043988
Publication Date: 2024-07-23
Author(s): Qiao Jin et al.
Primary Topic: Artificial Intelligence in Healthcare and Education

Overview

The study investigates the performance of Generative Pre-trained Transformer 4 with Vision (GPT-4V) in medical diagnostic tasks, specifically through its application in the New England Journal of Medicine (NEJM) Image Challenges. While previous research highlighted GPT-4V’s superiority over human physicians in multi-choice question accuracy, this study expands the evaluation to include the model’s rationales for image comprehension, recall of medical knowledge, and multimodal reasoning. The results indicate that GPT-4V achieves a multi-choice accuracy of 81.6%, surpassing the 77.8% accuracy of human physicians, and demonstrates over 78% accuracy in scenarios where physicians provide incorrect answers.

Despite these promising results, the study reveals significant concerns regarding the quality of GPT-4V’s rationales, particularly in image comprehension, where flawed reasoning was noted in 35.5% of correct final answers, with 27.2% specifically attributed to image comprehension errors. These findings underscore the importance of conducting more thorough evaluations of the rationales produced by GPT-4V before its potential integration into clinical settings, highlighting the need for caution in relying solely on its high accuracy in multi-choice assessments.

Methods

The “Methods” section outlines the experimental design and analytical techniques employed in the study. The researchers utilized a quantitative approach, employing statistical analyses to evaluate the data collected from various trials. Specific methodologies included the application of regression models to assess the relationship between the independent and dependent variables, ensuring robust results through appropriate controls and randomization.

Additionally, the study incorporated advanced computational techniques to analyze large datasets, allowing for the identification of significant patterns and correlations. The use of software tools for data visualization further facilitated the interpretation of results, enhancing the clarity of findings. Overall, the methodological framework established a solid foundation for the study’s conclusions, ensuring reliability and validity in the reported outcomes.

Discussion

In this section, the authors discuss the methodology and findings from their evaluation of the NEJM Image Challenge, which involved 207 recent medical imaging questions. These questions were categorized into three difficulty tiers based on user performance: “easy” (56-88% correct), “medium” (44-56%), and “hard” (28-44%). The authors employed GPT-4V and BiomedCLIP to assess the challenges, with GPT-4V prompted to analyze images and provide reasoning for its answers. BiomedCLIP, a multimodal language model, was tested in a zero-shot setting, achieving a low accuracy of 25.1%, indicating the significant challenge posed by these questions for smaller vision-language models.

The evaluation process included a two-stage annotation by medical professionals, who assessed GPT-4V’s responses across three categories: image comprehension, recall of medical knowledge, and step-by-step reasoning. Each response was rated as “Correct,” “Partially Correct,” or “Incorrect,” with detailed guidelines provided for each category. The inter-rater agreement among annotators was moderate to substantial, with Cohen Kappa values ranging from 0.470 to 0.698. Notably, GPT-4V outperformed or matched human physicians in six out of nine specialties, and there was a strong correlation between the ability to identify flaws in GPT-4V’s rationale and the accuracy of the physicians’ answers. This study uniquely combines quantitative performance metrics with systematic evaluations of the reasoning processes behind the model’s outputs.

شارك: