تقييم نماذج اللغة الكبيرة لاستشارة الآباء حول الساد الخلقي: السلامة، قابلية القراءة، وترجمة المعرفة للمعلومات التنموية والوراثية
Benchmarking large language models for congenital cataract parent counseling: safety, readability, and knowledge translation of developmental and genetic information

شارك:
المجلة: Frontiers in Cell and Developmental Biology، المجلد: 14
DOI: https://doi.org/10.3389/fcell.2026.1785731
PMID: https://pubmed.ncbi.nlm.nih.gov/41970961
تاريخ النشر: 2026-03-27
المؤلف: Liping Jiang وآخرون
الموضوع الرئيسي: علم الجينوم والأمراض النادرة

نظرة عامة

تقيّم هذه الدراسة أداء خمسة نماذج لغة كبيرة (LLMs) صينية متاحة للجمهور في تقديم الإرشادات حول الساد الخلقي (CC)، وهو سبب رئيسي لفقدان البصر لدى الأطفال. تم إجراء تحليل مقارن مقطعي باستخدام ثلاثين سؤالًا موحدًا تم تطويره بواسطة أطباء عيون كبار، مع التركيز على تطور العدسة، وعلم الأسباب، والمخاطر الجينية. تم تقييم النماذج من حيث كفاءة المخرجات، وبنية النص، وجودة المحتوى، حيث تم تقييم الردود من قبل أطباء عيون معصوبي الأعين على مقياس ليكرت من 5 نقاط عبر معايير مختلفة، بما في ذلك الدقة، والمنطق، والترابط، والسلامة، وسهولة الوصول إلى المحتوى.

كشفت النتائج عن اختلافات كبيرة في الأداء بين النماذج، حيث أظهر ChatGPT-5.2 أسرع وقت استجابة وميلاً نحو مخرجات منظمة، بينما أنتج DeepSeek-V3.1 أطول الردود. ومن الجدير بالذكر أن الدراسة وجدت أن الردود الأطول لم ترتبط بجودة أعلى؛ بل كانت العروض المنظمة مرتبطة بفهم أفضل وسلامة. تؤكد النتائج على إمكانيات نماذج اللغة الكبيرة لتعزيز التعليم حول الساد عند الأطفال واستشارات الأطباء، بينما تبرز أيضًا الحاجة إلى تحسين الموثوقية وقابلية التعميم من خلال التقييمات المستمرة ومعايير التحقق الموضوعية لضمان النشر الآمن والفعال في البيئات السريرية.

مقدمة

تسلط مقدمة ورقة البحث الضوء على الساد الخلقي (CC) كسبب رئيسي للعمى القابل للتجنب لدى الأطفال، بمعدل حدوث يقارب 1-15 لكل 10,000 طفل. إدارة CC أمر حاسم بسبب الطبيعة الحساسة للوقت في تطوير نظام الرؤية لدى الرضع، حيث يمكن أن تؤدي التأخيرات إلى حالات لا يمكن عكسها مثل الكسل أو الرأرأة. التشخيص المبكر، التدخل الجراحي في الوقت المناسب، والعلاج الشامل للكسل ضرورية لإعادة التأهيل البصري الفعال. ومع ذلك، غالبًا ما يواجه الآباء ضغوطًا نفسية وقلقًا خلال عمليات التشخيص والعلاج، مما يتفاقم بسبب نقص الوصول إلى المعلومات المتعلقة بالحالة وإدارتها.

تناقش الورقة إمكانيات نماذج اللغة الكبيرة (LLMs)، مثل ChatGPT وDeepSeek، لتحسين التواصل والتعليم للموظفين في طب العيون للأطفال. بينما يمكن أن توفر LLMs معلومات طبية قيمة، لا تزال هناك تحديات في ضمان دقة وسلامة المحتوى الذي تنتجه. تهدف الدراسة إلى تقييم منهجي لأداء خمسة نماذج LLMs مستخدمة على نطاق واسع في معالجة الأسئلة المتكررة من مقدمي الرعاية المتعلقة بـ CC، مع التركيز على جوانب رئيسية مثل تغطية المعلومات، والفهم، والتوصيات القابلة للتنفيذ. يسعى هذا التقييم إلى تقديم إرشادات قائمة على الأدلة للاستخدام الفعال لنماذج LLMs في تعزيز التعليم الصحي والاستشارة لمقدمي الرعاية للأطفال المصابين بـ CC.

النتائج

يقدم قسم “النتائج” في ورقة البحث النتائج الرئيسية المستمدة من التجارب أو التحليلات التي تم إجراؤها. يوضح النتائج الأساسية، بما في ذلك البيانات الإحصائية، والاتجاهات الملحوظة، وأي ارتباطات مهمة تم تحديدها. عادةً ما يتم توضيح النتائج من خلال الجداول، والرسوم البيانية، أو الأشكال، التي توفر تمثيلًا بصريًا للبيانات وتسهّل التفسير.

بالإضافة إلى ذلك، قد يناقش القسم تداعيات النتائج فيما يتعلق بفرضية البحث أو الأهداف. من الضروري ملاحظة أي نتائج غير متوقعة أو شذوذات ظهرت خلال الدراسة، حيث يمكن أن تقدم رؤى حول الآليات الأساسية أو تقترح مجالات لمزيد من التحقيق. بشكل عام، تسهم النتائج في الفهم الأوسع للموضوع وتدعم الاستنتاجات المستخلصة في الأقسام اللاحقة من الورقة.

المناقشة

تقدم هذه الدراسة تقييمًا مقطعيًا لجودة وسلامة الردود التي تولدها خمسة نماذج لغة كبيرة (LLMs) في معالجة أسئلة استشارة موجهة للآباء تتعلق بالساد الخلقي (CC). كانت أهداف البحث تقييم مدى فعالية هذه النماذج في التواصل بالمعلومات الحيوية المتعلقة بمرض العدسة النمائي وعلم الوراثة بطريقة آمنة ومفهومة لمقدمي الرعاية. شملت الدراسة عملية منظمة لتطوير مجموعة أسئلة تعكس مخاوف الآباء في العالم الحقيقي، تلتها تقييمات مستقلة لمخرجات النماذج عبر خمسة مجالات: الدقة، والمنطق، والترابط، والسلامة، وسهولة الوصول إلى المحتوى. تشير النتائج إلى أن ChatGPT-5.2 وGemini 3 Pro ظهرا كمعايير من حيث الدقة والسلامة، بينما لم ترتبط الردود الأطول بتقييمات جودة أعلى، مما يشير إلى أن طول الرد وحده ليس مؤشرًا موثوقًا لجودة المحتوى.

علاوة على ذلك، تبرز الدراسة أهمية المخرجات المنظمة، مثل التقسيم الواضح إلى فقرات واستخدام الجداول، والتي كانت مرتبطة بشكل إيجابي بالسلامة وسهولة الوصول. في المقابل، كانت أوقات الاستجابة مرتبطة سلبًا بمقاييس الجودة، مما يشير إلى أن الردود الأبطأ لم تعادل بالضرورة إجابات أفضل أو أكثر أمانًا. تؤكد النتائج على ضرورة التقييم المنظم والإشراف البشري في نشر نماذج LLMs للتعليم الصحي، خاصة في السياقات pediatrics حيث تكون قرارات مقدمي الرعاية حاسمة. يوصي المؤلفون بأن تعطي التطبيقات المستقبلية لنماذج LLMs في رعاية الساد عند الأطفال الأولوية لمعايير المخرجات المرتكزة على المخاطر وصيغ العرض المنظمة لتعزيز الوضوح والإرشادات القابلة للتنفيذ لمقدمي الرعاية.

Journal: Frontiers in Cell and Developmental Biology, Volume: 14
DOI: https://doi.org/10.3389/fcell.2026.1785731
PMID: https://pubmed.ncbi.nlm.nih.gov/41970961
Publication Date: 2026-03-27
Author(s): Liping Jiang et al.
Primary Topic: Genomics and Rare Diseases

Overview

This study evaluates the performance of five publicly accessible Chinese conversational large language models (LLMs) in providing guidance on congenital cataract (CC), a significant cause of childhood visual impairment. A cross-sectional comparative analysis was conducted using thirty standardized questions developed by senior ophthalmologists, focusing on lens development, etiology, and genetic risk. The models were assessed for output efficiency, textual structure, and content quality, with responses rated by blinded ophthalmologists on a 5-point Likert scale across various criteria, including Accuracy, Logic, Coherence, Safety, and Content Accessibility.

The results revealed significant differences in performance among the models, with ChatGPT-5.2 demonstrating the fastest response time and a tendency toward structured output, while DeepSeek-V3.1 produced the longest responses. Notably, the study found that longer responses did not correlate with higher quality; instead, structured presentations were linked to better comprehensibility and safety. The findings underscore the potential of LLMs to enhance pediatric cataract education and clinician counseling, while also highlighting the need for improved reliability and generalizability through ongoing assessments and objective verification metrics to ensure safe and effective deployment in clinical settings.

Introduction

The introduction of the research paper highlights congenital cataract (CC) as a significant cause of preventable childhood blindness, with an incidence rate of approximately 1-15 per 10,000 children. The management of CC is critical due to the time-sensitive nature of the infant visual system’s development, where delays can lead to irreversible conditions such as amblyopia or nystagmus. Early diagnosis, timely surgical intervention, and comprehensive amblyopia therapy are essential for effective visual rehabilitation. However, parents often face psychological distress and anxiety throughout the diagnostic and treatment processes, exacerbated by insufficient access to information regarding the condition and its management.

The paper discusses the potential of large language models (LLMs), such as ChatGPT and DeepSeek, to improve caregiver communication and education in pediatric ophthalmology. While LLMs can provide valuable medical information, challenges remain in ensuring the accuracy and safety of the content they generate. The study aims to systematically evaluate the performance of five widely used LLMs in addressing high-frequency caregiver questions related to CC, focusing on key aspects such as information coverage, comprehensibility, and actionable recommendations. This evaluation seeks to provide evidence-based guidance for the effective use of LLMs in enhancing health education and consultation for caregivers of children with CC.

Results

The “Results” section of the research paper presents key findings derived from the conducted experiments or analyses. It outlines the primary outcomes, including statistical data, observed trends, and any significant correlations identified. The results are typically illustrated through tables, graphs, or figures, which provide visual representation of the data and facilitate interpretation.

Additionally, the section may discuss the implications of the findings in relation to the research hypothesis or objectives. It is essential to note any unexpected results or anomalies that emerged during the study, as these can offer insights into the underlying mechanisms or suggest areas for further investigation. Overall, the results contribute to the broader understanding of the topic and support the conclusions drawn in subsequent sections of the paper.

Discussion

This study presents a cross-sectional evaluation of the quality and safety of responses generated by five large language models (LLMs) in addressing parent-oriented consultation questions related to congenital cataract (CC). The research aimed to assess how effectively these models communicate critical information regarding developmental lens pathology and genetics in a manner that is safe and comprehensible for caregivers. The study involved a structured process to develop a question set reflective of real-world parental concerns, followed by independent evaluations of model outputs across five domains: Accuracy, Logic, Coherence, Safety, and Content Accessibility. The findings indicate that ChatGPT-5.2 and Gemini 3 Pro emerged as benchmarks in terms of accuracy and safety, while longer responses did not correlate with higher quality ratings, suggesting that response length alone is not a reliable indicator of content quality.

Moreover, the study highlights the importance of structured outputs, such as clear segmentation into paragraphs and the use of tables, which were positively associated with safety and accessibility. In contrast, response times were negatively correlated with quality metrics, indicating that slower responses did not necessarily equate to better or safer answers. The results underscore the necessity for structured evaluation and human oversight in deploying LLMs for health education, particularly in pediatric contexts where caregiver decision-making is crucial. The authors recommend that future applications of LLMs in pediatric cataract care prioritize risk-centered output standards and structured presentation formats to enhance clarity and actionable guidance for caregivers.

شارك: