نماذج اللغة الكبيرة متعددة الأغراض تتفوق على أدوات الذكاء الاصطناعي السريرية المتخصصة في المعايير الطبية
General-purpose large language models outperform specialized clinical AI tools on medical benchmarks

شارك:
المجلة: Nature Medicine، المجلد: 32، العدد: 7
DOI: https://doi.org/10.1038/s41591-026-04431-5
PMID: https://pubmed.ncbi.nlm.nih.gov/42286322
تاريخ النشر: 2026-06-12
المؤلف: Krithik Vishwanath وآخرون
الموضوع الرئيسي: الذكاء الاصطناعي في الرعاية الصحية والتعليم

نظرة عامة

تقيّم هذه الفقرة أداء أدوات الذكاء الاصطناعي السريري المتخصصة، وتحديداً OpenEvidence و UpToDate Expert AI، مقابل النماذج اللغوية الكبيرة العامة الرائدة مثل GPT-5.2 و Gemini 3.1 Pro و Claude Opus 4.6. يتكون التقييم من ثلاث مراحل: (1) 500 سؤال MedQA تقيم المعرفة الطبية، (2) 500 عنصر HealthBench تقيس التوافق مع الأطباء الخبراء، و (3) معيار استفسارات سريرية حقيقية (RCQ) مستمد من 100 استفسار طبي غير محدد الهوية في بيئة سريرية. تضمنت مرحلة RCQ مراجعة عشوائية ومخفية من قبل 12 طبيباً أمريكياً، مما أسفر عن 1,800 تعليق.

تشير النتائج إلى أن النماذج اللغوية الكبيرة الرائدة تفوقت باستمرار على أدوات الذكاء الاصطناعي السريري عبر جميع مراحل التقييم. على وجه التحديد، حقق Gemini أعلى دقة بنسبة 97.4% في أسئلة MedQA، تلاه GPT بنسبة 94.2% و Claude بنسبة 90.2%. في المقابل، سجلت OpenEvidence و UpToDate 89.6% و 88.4%، على التوالي. كشفت التحليلات الإحصائية عن فروق أداء كبيرة، حيث تفوق Gemini على جميع النماذج الأخرى (McNemar $P < 1 \times 10^{-4}$ مقابل OpenEvidence و UpToDate و Claude؛ $P = 0.02$ مقابل GPT). تؤكد هذه النتائج على ضرورة إجراء تقييمات مستقلة وعالمية لأدوات الذكاء الاصطناعي قبل دمجها في الممارسة السريرية، حيث تفتقر النماذج الحالية المملوكة إلى التحقق الشفاف.

الطرق

تحدد فقرة “الطرق” في ورقة البحث التصميم التجريبي والتقنيات التحليلية المستخدمة للتحقيق في سؤال البحث. استخدمت الدراسة نهجاً كمياً، يتضمن تحليلات إحصائية لتقييم البيانات المجمعة من تجارب مختلفة. تضمنت المنهجيات المحددة تجارب مختبرية محكومة، حيث تم التلاعب بالمتغيرات بشكل منهجي لملاحظة تأثيراتها على النتائج ذات الصلة.

شملت جمع البيانات استخدام أدوات موحدة لضمان الموثوقية والصلاحية. تم إجراء التحليل باستخدام أدوات برمجية سهلت تطبيق الاختبارات الإحصائية المناسبة، مثل اختبارات t و ANOVA، لتحديد الفروق المهمة بين المجموعات. كما تتناول الفقرة استراتيجية العينة، بما في ذلك معايير اختيار المشاركين وحسابات حجم العينة التي تم إجراؤها لضمان قوة كافية للتحليلات الإحصائية. بشكل عام، كانت الطرق المستخدمة مصممة بدقة لتحقيق نتائج قوية وقابلة للتكرار.

المناقشة

تقدم فقرة المناقشة في ورقة البحث تحليلاً مقارناً للنماذج اللغوية الكبيرة الرائدة (LLMs) وأدوات الذكاء الاصطناعي السريري في الاستجابة لاستفسارات الأطباء في العالم الحقيقي. تشير النتائج إلى أن النماذج الرائدة، وتحديداً Gemini و GPT-5.2 و Claude، تفوقت بشكل كبير على الأدوات السريرية مثل OpenEvidence و UpToDate عبر أبعاد مختلفة، بما في ذلك الدقة السريرية، والكمال، والسلامة، والوضوح. كانت احتمالات حصول أدوات الذكاء الاصطناعي السريري على تقييمات أعلى من Gemini أقل بكثير، مع نسب احتمالات تتراوح من 0.13 إلى 0.51 (جميعها $P < 0.0001$). كما أبرزت الدراسة أن Google AI Overview أدت أداءً مماثلاً للأدوات السريرية، ومع ذلك حافظت جميع النماذج على نتائج سلامة مماثلة، دون وجود فروق كبيرة في المحتوى الضار أو الهلوسات. تشير النتائج إلى أن الأداء المتفوق للنماذج الرائدة قد ينبع من قدراتها المحسنة في استرجاع المعرفة والتفكير، فضلاً عن دورات تطويرها السريعة ومجموعات بيانات التدريب الأكبر. تعترف الدراسة بالقيود، مثل طريقة الاستفسار لأدوات الذكاء الاصطناعي السريري والتحيزات المحتملة في تقييمات المعايير. وتؤكد على الحاجة إلى أطر تقييم مستقلة تعكس البيئات السريرية الحقيقية، داعية إلى دمج نماذج لغوية كبيرة محددة للمستشفيات لتحسين الفائدة السريرية مع ضمان السلامة. بشكل عام، تؤكد النتائج على المشهد المتطور للذكاء الاصطناعي في الرعاية الصحية، مشيرة إلى أنه بينما قد تمتلك الأدوات السريرية شرعية مؤسسية، إلا أنها لا تتفوق على النماذج الرائدة في المعرفة وفعالية التواصل.

Journal: Nature Medicine, Volume: 32, Issue: 7
DOI: https://doi.org/10.1038/s41591-026-04431-5
PMID: https://pubmed.ncbi.nlm.nih.gov/42286322
Publication Date: 2026-06-12
Author(s): Krithik Vishwanath et al.
Primary Topic: Artificial Intelligence in Healthcare and Education

Overview

This section evaluates the performance of specialized clinical artificial intelligence (AI) tools, specifically OpenEvidence and UpToDate Expert AI, against leading general-purpose large language models (LLMs) such as GPT-5.2, Gemini 3.1 Pro, and Claude Opus 4.6. The evaluation consists of three stages: (1) 500 MedQA questions assessing medical knowledge, (2) 500 HealthBench items measuring alignment with expert clinicians, and (3) a real clinical queries (RCQ) benchmark derived from 100 de-identified physician queries in a clinical setting. The RCQ stage involved a randomized, blinded review by 12 US clinicians, resulting in 1,800 annotations.

The findings indicate that frontier LLMs consistently outperformed the clinical AI tools across all evaluation stages. Specifically, Gemini achieved the highest accuracy of 97.4% on the MedQA questions, followed by GPT at 94.2% and Claude at 90.2%. In contrast, OpenEvidence and UpToDate scored 89.6% and 88.4%, respectively. Statistical analysis revealed significant performance differences, with Gemini outperforming all other models (McNemar $P < 1 \times 10^{-4}$ versus OpenEvidence, UpToDate, and Claude; $P = 0.02$ versus GPT). These results underscore the necessity for independent, real-world evaluations of AI tools prior to their integration into clinical practice, as current proprietary models lack transparent validation.

Methods

The “Methods” section of the research paper outlines the experimental design and analytical techniques employed to investigate the research question. The study utilized a quantitative approach, incorporating statistical analyses to evaluate the data collected from various experiments. Specific methodologies included controlled laboratory experiments, where variables were systematically manipulated to observe their effects on the outcomes of interest.

Data collection involved the use of standardized instruments to ensure reliability and validity. The analysis was conducted using software tools that facilitated the application of appropriate statistical tests, such as t-tests and ANOVA, to determine significant differences between groups. The section also details the sampling strategy, including the selection criteria for participants and the sample size calculations that were performed to ensure adequate power for the statistical analyses. Overall, the methods employed were rigorously designed to yield robust and replicable results.

Discussion

The discussion section of the research paper presents a comparative analysis of frontier large language models (LLMs) and clinical AI tools in responding to real-world physician queries. The findings indicate that frontier models, specifically Gemini, GPT-5.2, and Claude, significantly outperformed clinical tools like OpenEvidence and UpToDate across various dimensions, including clinical correctness, completeness, safety, and clarity. The odds of clinical AI tools receiving higher ratings than Gemini were notably lower, with odds ratios ranging from 0.13 to 0.51 (all $P < 0.0001$). The study also highlighted that Google AI Overview performed comparably to clinical tools, yet all models maintained similar safety outcomes, with no significant differences in harmful content or hallucinations. The results suggest that the superior performance of frontier models may stem from their enhanced knowledge retrieval and reasoning capabilities, as well as their rapid development cycles and larger training datasets. The study acknowledges limitations, such as the querying method for clinical tools and potential biases in benchmark evaluations. It emphasizes the need for independent evaluation frameworks that reflect real clinical environments, advocating for the integration of hospital-specific LLMs to optimize clinical utility while ensuring safety. Overall, the findings underscore the evolving landscape of AI in healthcare, indicating that while clinical tools may possess institutional legitimacy, they do not surpass frontier models in knowledge and communication effectiveness.

شارك: