ChatGPT-5 مقابل نماذج اللغة الكبيرة السائدة الأخرى في استفسارات مرضى اعتلال الشبكية السكري الأساسية
ChatGPT-5 versus other mainstream large language models in core diabetic retinopathy patient queries

شارك:
المجلة: Frontiers in Cell and Developmental Biology، المجلد: 14
DOI: https://doi.org/10.3389/fcell.2026.1754221
PMID: https://pubmed.ncbi.nlm.nih.gov/41960186
تاريخ النشر: 2026-03-25
المؤلف: Xiaomin Cang وآخرون
الموضوع الرئيسي: الذكاء الاصطناعي في الرعاية الصحية والتعليم

نظرة عامة

تبحث الدراسة في أداء خمسة نماذج لغوية كبيرة (LLMs) في تقديم معلومات تتعلق باعتلال الشبكية السكري (DR)، وهو سبب رئيسي لفقدان الرؤية القابل للتجنب. تم تصنيف مجموعة من ثلاثين سؤالًا من المرضى إلى خمسة مجالات، وتم تقييم الردود بناءً على الكفاءة وجودة المحتوى والخصائص الهيكلية. وجدت الدراسة اختلافات كبيرة بين النماذج، حيث تفوق ChatGPT-5 على الآخرين في سرعة الاستجابة (15.92 ثانية) وجودة المحتوى، محققًا أعلى درجات متوسطة للدقة (5.00) والتناسق المنطقي (4.50). في المقابل، أظهر Kimi أدنى دقة (3.50)، بينما أنتج DeepSeek-V3.1 أطول الردود.

تشير النتائج إلى أن فعالية LLMs في استشارات المرضى تعتمد على النموذج، مما يبرز أن الردود الأقصر والأكثر وضوحًا التي تعالج النقاط الرئيسية بدقة تكون أكثر قيمة من الإجابات الأطول والأقل دقة. تختتم الدراسة بالقول إنه بينما يمكن أن تكون LLMs أدوات مفيدة لتعليم DR والفرز الأولي، يجب ألا تحل محل اتخاذ القرارات السريرية، خاصة في الحالات الحادة. تشمل التوصيات لتحسين النماذج المطورة في الصين التركيز على التماسك المنطقي والتواصل الموجز لتعزيز قابلية الاستخدام في السياقات السريرية.

مقدمة

اعتلال الشبكية السكري (DR) هو مضاعفة دقيقة شائعة لمرض السكري، تسهم بشكل كبير في ضعف الرؤية بين البالغين في سن العمل على مستوى العالم. تشير التقديرات الأخيرة إلى أن حوالي 22.3% من الأفراد المصابين بالسكري يتأثرون بـ DR، مع 6.2% يعانون من أشكال تهدد الرؤية. من المتوقع أن ترتفع انتشار DR العالمي من 103 مليون في 2020 إلى 160 مليون بحلول 2045، مما يزيد من تفاقمها بسبب ارتفاع معدلات السكري وشيخوخة السكان. على الرغم من فعالية الفحص المبكر وتعليم المرضى في تقليل فقدان الرؤية، لا يزال التغطية العالمية غير كافية. تقدم التطورات الأخيرة في الذكاء الاصطناعي (AI)، وخاصة النماذج اللغوية الكبيرة (LLMs)، فرصًا جديدة لتعزيز نشر المعلومات الطبية وتعليم المرضى في سياق DR.

تقيّم هذه الدراسة أداء عدة LLMs في تقديم تعليم حول المرض واستشارات المرضى لـ DR. من خلال استخدام إطار تقييم خبراء كمي متعدد الأبعاد، تهدف الدراسة إلى تقييم جدوى وسلامة LLMs في إدارة أمراض الشبكية. تشير النتائج إلى وجود اختلافات كبيرة في كفاءة المخرجات وجودة المحتوى بين خمسة LLMs رئيسية. بشكل ملحوظ، أظهر ChatGPT-5 أوقات استجابة وجودة محتوى متفوقة عبر عدة أبعاد، بما في ذلك الدقة والمنطق والتماسك والسلامة وسهولة الوصول إلى المحتوى، مقارنة بالنماذج الأخرى. تؤكد هذه النتائج على إمكانيات LLMs، وخاصة ChatGPT-5، في دعم تعليم المرضى المدعوم بالذكاء الاصطناعي والتواصل الصحي العام بشأن DR.

النتائج

يقدم قسم “النتائج” من ورقة البحث النتائج الرئيسية المستمدة من التجارب أو التحليلات التي تم إجراؤها. يوضح النتائج التي توصلت إليها الدراسة، مع تسليط الضوء على الاتجاهات البيانية المهمة، والتحليلات الإحصائية، وأي ارتباطات أو أنماط تم ملاحظتها. غالبًا ما تكون النتائج مصحوبة بأرقام وجداول ذات صلة توضح البيانات بصريًا، مما يعزز فهم القارئ للنتائج.

قد يناقش القسم أيضًا تداعيات هذه النتائج فيما يتعلق بالفرضيات المطروحة في بداية البحث. من المهم ملاحظة أي نتائج غير متوقعة أو شذوذ ظهرت خلال الدراسة، حيث يمكن أن توفر هذه رؤى حول مجالات لمزيد من التحقيق. بشكل عام، تعتبر النتائج أساسًا للنقاشات اللاحقة والاستنتاجات المستخلصة في الورقة.

المناقشة

استخدمت الدراسة تصميمًا مقارنًا مقطعيًا لتقييم أداء خمسة نماذج لغوية كبيرة (LLMs) في الرد على أسئلة اعتلال الشبكية السكري (DR) التي بدأها المرضى. تم تطوير مجموعة موحدة من 30 سؤالًا من خلال توافق الخبراء، وتم استجواب كل نموذج في ظل ظروف متطابقة لضمان القابلية للمقارنة. شملت النماذج التي تم تقييمها ChatGPT-5 وDeepSeek-V3.1 وDoubao وWenxinyiyan 4.5 Turbo وKimi. أشارت النتائج إلى أن ChatGPT-5 تفوق على الآخرين في الدقة والمنطق وسهولة الوصول إلى المحتوى، محققًا درجات متوسطة تبلغ 5.0 و4.5 في هذه الأبعاد. في المقابل، احتل Kimi المرتبة الأدنى، خاصة في الدقة، مما يبرز التباين الكبير في الأداء بين النماذج.

تشير النتائج إلى أنه بينما تعزز بنية ChatGPT-5 المتقدمة، التي تتضمن نظام توجيه موحد ثنائي الوضع، قدراته على التفكير وتقلل من الهلاوس، قد تعطي نماذج أخرى مثل DeepSeek-V3.1 وKimi الأولوية للبلاغة على الدقة الواقعية في إعدادات عدم وجود بيانات. من الجدير بالذكر أن الدراسة كشفت أن LLMs الصينية، على الرغم من مزاياها اللغوية، لم تحقق دقة أو تفكير منطقي قابل للمقارنة، مما يشير إلى الحاجة لمزيد من التوافق في المعرفة الطبية. تؤكد النتائج على إمكانيات LLMs كأدوات مساعدة للتعليم الصحي والفرز الأولي في إدارة DR، مع التأكيد على أهمية التواصل الواضح وتحديد المخاطر في تفاعلات المرضى. تشمل قيود الدراسة تنسيق الاستفسار ذو الدور الواحد وغياب ديناميات الاستشارة في العالم الحقيقي، مما يشير إلى أن مزيدًا من البحث ضروري للتحقق من هذه النتائج عبر سياقات ونماذج متنوعة.

Journal: Frontiers in Cell and Developmental Biology, Volume: 14
DOI: https://doi.org/10.3389/fcell.2026.1754221
PMID: https://pubmed.ncbi.nlm.nih.gov/41960186
Publication Date: 2026-03-25
Author(s): Xiaomin Cang et al.
Primary Topic: Artificial Intelligence in Healthcare and Education

Overview

The research investigates the performance of five large language models (LLMs) in providing information related to diabetic retinopathy (DR), a major cause of preventable vision loss. A set of thirty patient questions was categorized into five domains, and responses were evaluated based on efficiency, content quality, and structural characteristics. The study found significant differences among the models, with ChatGPT-5 outperforming others in response speed (15.92 seconds) and content quality, achieving the highest median scores for accuracy (5.00) and logical consistency (4.50). In contrast, Kimi exhibited the lowest accuracy (3.50), while DeepSeek-V3.1 produced the longest responses.

The findings suggest that the effectiveness of LLMs in patient consultations is model-dependent, emphasizing that shorter, clearer responses that accurately address key points are more valuable than longer, less precise answers. The study concludes that while LLMs can serve as useful tools for DR education and preliminary triage, they should not replace clinical decision-making, especially in acute situations. Recommendations for optimizing Chinese-developed models include focusing on logical coherence and concise communication to enhance usability in clinical contexts.

Introduction

Diabetic retinopathy (DR) is a prevalent microvascular complication of diabetes, significantly contributing to vision impairment among working-age adults globally. Recent estimates indicate that approximately 22.3% of individuals with diabetes are affected by DR, with 6.2% experiencing vision-threatening forms. The global prevalence of DR is projected to rise from 103 million in 2020 to 160 million by 2045, exacerbated by increasing diabetes rates and an aging population. Despite the effectiveness of early screening and patient education in reducing vision loss, global coverage remains inadequate. Recent advancements in artificial intelligence (AI), particularly large language models (LLMs), present new opportunities for enhancing medical information dissemination and patient education in the context of DR.

This study evaluates the performance of multiple LLMs in providing disease education and patient consultation for DR. By employing a multidimensional quantitative expert rating framework, the research aims to assess the feasibility and safety of LLMs in retinal disease management. The findings indicate significant differences in output efficiency and content quality among five mainstream LLMs. Notably, ChatGPT-5 demonstrated superior response times and content quality across multiple dimensions, including accuracy, logic, coherence, safety, and content accessibility, compared to other models. These results underscore the potential of LLMs, particularly ChatGPT-5, in supporting AI-assisted patient education and public health communication regarding DR.

Results

The “Results” section of the research paper presents the key findings derived from the conducted experiments or analyses. It details the outcomes of the study, highlighting significant data trends, statistical analyses, and any observed correlations or patterns. The results are often accompanied by relevant figures and tables that illustrate the data visually, enhancing the reader’s understanding of the findings.

The section may also discuss the implications of these results in relation to the hypotheses posed at the outset of the research. It is crucial to note any unexpected outcomes or anomalies that emerged during the study, as these can provide insights into areas for further investigation. Overall, the results serve as a foundation for the subsequent discussion and conclusions drawn in the paper.

Discussion

The study employed a cross-sectional comparative design to evaluate the performance of five mainstream large language models (LLMs) in responding to patient-initiated diabetic retinopathy (DR) questions. A standardized set of 30 questions was developed through expert consensus, and each model was queried under identical conditions to ensure comparability. The models assessed included ChatGPT-5, DeepSeek-V3.1, Doubao, Wenxinyiyan 4.5 Turbo, and Kimi. Results indicated that ChatGPT-5 outperformed the others in accuracy, logic, and content accessibility, achieving median scores of 5.0 and 4.5 in these dimensions. In contrast, Kimi ranked lowest, particularly in accuracy, highlighting significant performance variability among the models.

The findings suggest that while ChatGPT-5’s advanced architecture, which incorporates a unified-routing dual-mode system, enhances its reasoning capabilities and reduces hallucinations, other models like DeepSeek-V3.1 and Kimi may prioritize fluency over factual accuracy in zero-shot settings. Notably, the study revealed that the Chinese LLMs, despite their linguistic advantages, did not achieve comparable accuracy or logical reasoning, indicating a need for further alignment in medical knowledge. The results underscore the potential of LLMs as auxiliary tools for health education and preliminary triage in DR management, while emphasizing the importance of clear communication and risk delineation in patient interactions. Limitations of the study include the single-turn query format and the lack of real-world consultation dynamics, suggesting that further research is necessary to validate these findings across diverse contexts and model iterations.

شارك: