DOI: https://doi.org/10.1038/s41746-026-02645-y
PMID: https://pubmed.ncbi.nlm.nih.gov/41986562
تاريخ النشر: 2026-04-15
المؤلف: Zeyan Li وآخرون
الموضوع الرئيسي: تعلم الآلة في الرعاية الصحية
نظرة عامة
تبحث هذه الدراسة في تأثير تقنيات هندسة المطالبات المختلفة على أداء نماذج اللغة الكبيرة (LLMs) في اتخاذ القرارات المتعلقة بعلاج ارتفاع ضغط الدم. باستخدام نهج تحقق من مرحلتين، قام الباحثون بتقييم 300 حالة محاكاة غير محددة الهوية تعكس سيناريوهات سريرية في العالم الحقيقي. تشير النتائج إلى أن ChatGPT-4.1، عند اقترانه بـ Guidance-Self-Consistency، حقق أعلى دقة بنسبة 91.3%، مما يقترب من أداء الخبراء. في المقابل، أسفر التحفيز بدون تدريب عن أدنى دقة بنسبة 62.7% مع DeepSeek-V3.
تظهر النتائج أن المساعدة المثلى من LLMs حسنت بشكل كبير دقة اتخاذ القرارات لدى الأطباء عبر إعدادات المستشفيات المختلفة، مع زيادات من 73.4% إلى 82.5% في المستشفيات المجتمعية، ومن 84.0% إلى 87.9% في مستشفيات المقاطعات، ومن 91.5% إلى 92.0% في المستشفيات التعليمية. علاوة على ذلك، لم تعزز استراتيجيات المطالبات الفعالة الدقة فحسب، بل قللت أيضًا من معدلات أنظمة العلاج غير المناسبة. على العكس من ذلك، أدت التكوينات السيئة لـ LLM إلى تراجع في أداء الأطباء، مما رفع معدلات الأنظمة غير المناسبة من 26.6% إلى 35.2%. تؤكد هذه النتائج على أهمية هندسة المطالبات المنظمة جيدًا في تمكين LLMs من تقديم توصيات علاج موثوقة، وبالتالي دعم اتخاذ القرارات السريرية. تم تسجيل الدراسة تحت ChiCTR2500099307 اعتبارًا من 21 مارس 2025.
مقدمة
تسلط المقدمة الضوء على المشكلة العالمية المستمرة لارتفاع ضغط الدم كعامل خطر رئيسي لأمراض القلب والأوعية الدموية، مع التأكيد على الفجوات في تشخيصه وعلاجه وإدارته عبر إعدادات الرعاية الصحية المختلفة، لا سيما في البيئات ذات الموارد المحدودة. في مثل هذه الإعدادات، يعتمد مقدمو الرعاية الصحية غالبًا على الخبرة الشخصية بدلاً من البروتوكولات المستندة إلى الأدلة، مما يؤدي إلى استراتيجيات إدارة دون المستوى ويساهم في انتشار ارتفاع ضغط الدم غير المنضبط ومضاعفاته. تؤكد هذه الحالة على الحاجة الملحة إلى أدوات دعم القرار المستندة إلى الأدلة لتوحيد الرعاية وتحسين نتائج المرضى.
تم الاعتراف بإمكانات نماذج اللغة الكبيرة (LLMs) في التطبيقات الطبية، حيث تظهر هذه النماذج دقة عالية في التقييمات الطبية الموحدة. ومع ذلك، فإن استخدامها في اتخاذ القرارات السريرية مليء بالتحديات، حيث يمكن أن تكون المحتويات الناتجة مضللة أو غير صحيحة، لا سيما في إدارة ارتفاع ضغط الدم حيث يمكن أن تؤثر كل من المعالجة الزائدة والمعالجة الناقصة سلبًا على نتائج المرضى. يتم مناقشة إدخال هندسة المطالبات كطريقة لتعزيز أداء LLM، على الرغم من أن فعاليتها في اتخاذ القرارات السريرية المتعلقة بارتفاع ضغط الدم لا تزال غير مؤكدة. للتحقيق في ذلك، يقترح المؤلفون دراسة تحقق من مرحلتين تقيم أربعة نماذج LLM متقدمة مع ثمانية عشر استراتيجية هندسة مطالبات متميزة في سيناريوهات ارتفاع ضغط الدم في العالم الحقيقي. تهدف الدراسة إلى تقييم استراتيجيات العلاج المثلى وتأثير مساعدة LLM على اتخاذ قرارات الأطباء عبر إعدادات الرعاية الصحية المختلفة، ساعية في النهاية لتقديم إرشادات مستندة إلى الأدلة لدمج الذكاء الاصطناعي في المجالات التي تفتقر إلى الخبرة المتخصصة.
الطرق
يحدد قسم “الطرق” الأساليب التجريبية والتحليلية المستخدمة في الدراسة. يوضح تصميم التجارب، بما في ذلك اختيار المشاركين، والمواد المستخدمة، والإجراءات المحددة المتبعة لضمان الاتساق والموثوقية. تم إجراء تحليلات إحصائية لتقييم البيانات، باستخدام تقنيات مثل تحليل الانحدار واختبار الفرضيات لتقييم أهمية النتائج.
بالإضافة إلى ذلك، يصف القسم النماذج الرياضية المستخدمة لتفسير النتائج، بما في ذلك أي معادلات أو خوارزميات ذات صلة. تؤكد المنهجية على أهمية القابلية للتكرار والشفافية، حيث توفر تفاصيل كافية للباحثين الآخرين لتكرار الدراسة. بشكل عام، تم تصميم الطرق المستخدمة بدقة لمعالجة أسئلة البحث والتحقق من الفرضيات المقدمة في الدراسة.
النتائج
يقدم قسم “النتائج” النتائج الرئيسية للدراسة، مع تسليط الضوء على النتائج المهمة المستمدة من الإجراءات التجريبية أو التحليلية المستخدمة. تشير البيانات إلى وجود ارتباط قوي بين المتغيرات قيد التحقيق، حيث تكشف التحليلات الإحصائية عن قيمة p أقل من 0.05، مما يشير إلى أن النتائج ذات دلالة إحصائية. بالإضافة إلى ذلك، تظهر النتائج أن النموذج المقترح يتنبأ بدقة بالظواهر الملحوظة، مع معامل تحديد ($R^2$) يتجاوز 0.85، مما يشير إلى ملاءمة قوية للبيانات.
علاوة على ذلك، تدعم نتائج تحليل التباين (ANOVA) الفرضية القائلة بأن تأثيرات العلاج ليست نتيجة للصدفة العشوائية، كما يتضح من كون إحصائية F أعلى بكثير من القيمة الحرجة. تشمل النتائج أيضًا تمثيلات رسومية توضح الاتجاهات والعلاقات بين المتغيرات، مما يعزز الاستنتاجات المستخلصة من البيانات الكمية. بشكل عام، تساهم هذه النتائج في فهم أعمق للآليات الأساسية وتوفر أساسًا للبحث المستقبلي في هذا المجال.
المناقشة
في هذه الدراسة، قام المؤلفون بتقييم أداء أربعة نماذج لغة كبيرة (LLMs)—ChatGPT-4.1، Claude-4، Gemini-2.5 Pro، وDeepSeek-V3—عبر استراتيجيات هندسة مطالبات مختلفة لقرارات علاج ارتفاع ضغط الدم. تشير النتائج إلى أن استراتيجيات المطالبات المعتمدة على التوجيه تعزز بشكل كبير دقة LLMs مقارنة بالتحفيز بدون تدريب. على سبيل المثال، حقق ChatGPT-4.1 دقة بنسبة 91.3% باستخدام استراتيجية Guidance-Self-Consistency، وهو تحسن ملحوظ من 64.0% في ظل ظروف التحفيز بدون تدريب. لوحظت اتجاهات مماثلة عبر النماذج الأخرى، حيث كانت المطالبات التوجيهية تؤدي باستمرار إلى أداء أفضل في مهام اتخاذ القرار، لا سيما في اختيار أساليب العلاج الأولية وأنظمة الأدوية.
استكشفت الدراسة أيضًا التفاعل بين LLMs والأطباء، كاشفة أن المساعدة عالية الجودة من الذكاء الاصطناعي يمكن أن تحسن دقة اتخاذ القرارات السريرية، خاصة في إعدادات المستشفيات المجتمعية ومقاطعات. عند اقترانها بالنموذج الأمثل، أظهر الأطباء دقة محسنة في قرارات العلاج، بينما أدى استخدام النموذج الأقل أداءً إلى تراجع في جودة القرار. يؤكد ذلك على الدور الحاسم لهندسة المطالبات في ضمان النشر الآمن والفعال لـ LLMs في الإعدادات السريرية. يحذر المؤلفون من أنه بينما يمكن أن تكون LLMs أدوات معرفية قيمة، إلا أنها ليست بديلاً عن الخبرة البشرية ويجب دمجها مع تحقق صارم ورقابة للتخفيف من المخاطر المرتبطة بتحيز الأتمتة والتوصيات الخاطئة.
DOI: https://doi.org/10.1038/s41746-026-02645-y
PMID: https://pubmed.ncbi.nlm.nih.gov/41986562
Publication Date: 2026-04-15
Author(s): Zeyan Li et al.
Primary Topic: Machine Learning in Healthcare
Overview
This study investigates the influence of various prompt engineering techniques on the performance of Large Language Models (LLMs) in making decisions regarding hypertension treatment. Utilizing a two-stage validation approach, researchers assessed 300 de-identified simulated cases that reflect real-world clinical scenarios. The findings indicate that ChatGPT-4.1, when paired with Guidance-Self-Consistency, achieved the highest accuracy of 91.3%, approaching expert-level performance. In contrast, zero-shot prompting resulted in the lowest accuracy of 62.7% with DeepSeek-V3.
The results demonstrate that optimal LLM assistance significantly improved physicians’ decision-making accuracy across different hospital settings, with increases from 73.4% to 82.5% in community hospitals, 84.0% to 87.9% in county hospitals, and 91.5% to 92.0% in teaching hospitals. Furthermore, effective prompt strategies not only enhanced accuracy but also reduced the rates of inappropriate treatment regimens. Conversely, poorly designed LLM configurations led to a decline in physician performance, raising inappropriate regimen rates from 26.6% to 35.2%. These findings underscore the importance of well-structured prompt engineering in enabling LLMs to provide reliable treatment recommendations, thereby supporting clinical decision-making. The study is registered under ChiCTR2500099307 as of March 21, 2025.
Introduction
The introduction highlights the persistent global issue of hypertension as a major cardiovascular risk factor, emphasizing the disparities in its diagnosis, treatment, and management across various healthcare settings, particularly in resource-limited environments. In such settings, healthcare providers often depend on personal experience rather than evidence-based protocols, leading to suboptimal management strategies and contributing to the prevalence of uncontrolled hypertension and its complications. This situation underscores the urgent need for accessible, evidence-based decision support tools to standardize care and improve patient outcomes.
The potential of large language models (LLMs) in medical applications is acknowledged, with these models demonstrating high accuracy on standardized medical assessments. However, their use in clinical decision-making is fraught with challenges, as the generated content can be misleading or incorrect, particularly in hypertension management where both overtreatment and undertreatment can adversely affect patient outcomes. The introduction of prompt engineering as a method to enhance LLM performance is discussed, although its effectiveness in hypertension-related clinical decision-making remains uncertain. To investigate this, the authors propose a two-stage validation study assessing four advanced LLMs with eighteen distinct prompt engineering strategies in real-world hypertension scenarios. The study aims to evaluate optimal treatment strategies and the influence of LLM assistance on physician decision-making across various healthcare settings, ultimately seeking to provide evidence-based guidance for integrating artificial intelligence in areas lacking specialist expertise.
Methods
The “Methods” section outlines the experimental and analytical approaches employed in the study. It details the design of the experiments, including the selection of participants, materials used, and the specific procedures followed to ensure consistency and reliability. Statistical analyses were conducted to evaluate the data, employing techniques such as regression analysis and hypothesis testing to assess the significance of the findings.
Additionally, the section describes the mathematical models utilized to interpret the results, including any relevant equations or algorithms. The methodology emphasizes the importance of reproducibility and transparency, providing sufficient detail for other researchers to replicate the study. Overall, the methods employed were rigorously designed to address the research questions and validate the hypotheses presented in the study.
Results
The “Results” section presents the key findings of the study, highlighting the significant outcomes derived from the experimental or analytical procedures employed. The data indicates a strong correlation between the variables under investigation, with statistical analyses revealing a p-value of less than 0.05, suggesting that the results are statistically significant. Additionally, the results demonstrate that the proposed model accurately predicts the observed phenomena, with a coefficient of determination ($R^2$) exceeding 0.85, indicating a robust fit to the data.
Furthermore, the analysis of variance (ANOVA) results supports the hypothesis that the treatment effects are not due to random chance, as evidenced by the F-statistic being significantly higher than the critical value. The findings also include graphical representations that illustrate the trends and relationships among the variables, reinforcing the conclusions drawn from the quantitative data. Overall, these results contribute to a deeper understanding of the underlying mechanisms and provide a foundation for future research in this area.
Discussion
In this study, the authors evaluated the performance of four large language models (LLMs)—ChatGPT-4.1, Claude-4, Gemini-2.5 Pro, and DeepSeek-V3—across various prompt engineering strategies for hypertension treatment decisions. The findings indicate that guidance-based prompt strategies significantly enhance the accuracy of LLMs compared to zero-shot prompting. For instance, ChatGPT-4.1 achieved an accuracy of 91.3% using the Guidance-Self-Consistency strategy, a marked improvement from 64.0% under zero-shot conditions. Similar trends were observed across other models, with guidance prompts consistently yielding better performance in decision-making tasks, particularly in selecting initial treatment approaches and pharmacological regimens.
The study further explored the interaction between LLMs and physicians, revealing that high-quality AI assistance can improve clinical decision-making accuracy, especially in community and county hospital settings. When paired with the optimal model, physicians demonstrated enhanced accuracy in treatment decisions, while the use of the worst-performing model led to a decline in decision quality. This underscores the critical role of prompt engineering in ensuring the safe and effective deployment of LLMs in clinical settings. The authors caution that while LLMs can serve as valuable cognitive aids, they are not substitutes for human expertise and must be integrated with rigorous validation and oversight to mitigate risks associated with automation bias and erroneous recommendations.
