DOI: https://doi.org/10.3389/fpubh.2026.1777577
PMID: https://pubmed.ncbi.nlm.nih.gov/41799487
تاريخ النشر: 2026-02-20
المؤلف: Yanping Zhang وآخرون
الموضوع الرئيسي: الذكاء الاصطناعي في الرعاية الصحية والتعليم
نظرة عامة
في هذه الدراسة، يحقق المؤلفون في جودة وقابلية قراءة المعلومات الصحية التي تم إنشاؤها بواسطة خمسة نماذج لغوية كبيرة (LLMs)—دو باو، ديب سيك، وين شين يي يان، تونغ يي تشيان وين، وGPT-5—باستخدام مجموعة موحدة من 20 سؤالاً متكرراً حول سرطان الجلد باللغة الصينية الماندرين. استخدمت التقييمات أداة تقييم مواد التعليم للمرضى الصديقة للمستهلك (c-PEMAT-P) ومقياس الجودة العالمي (GQS) لتقييم جودة المحتوى، بينما تم قياس قابلية القراءة باستخدام سبعة مؤشرات مختلفة، بما في ذلك مؤشر قابلية القراءة الآلي (ARI) ودرجة سهولة القراءة فليش (FRES).
تظهر النتائج أنه بينما أظهرت النماذج مستويات مشابهة من الفهم وقابلية التنفيذ وفقًا لـ c-PEMAT-P، كانت هناك اختلافات كبيرة في الجودة العامة، حيث حقق GPT-5 أعلى درجة في GQS. تفاوتت قابلية القراءة بشكل كبير عبر النماذج وفئات المحتوى، مما يشير إلى أنه لا يوجد نموذج واحد يتفوق في جميع مقاييس قابلية القراءة. من الجدير بالذكر أن تحليلات الارتباط أظهرت أن الجودة وقابلية القراءة مستقلة إلى حد كبير، مما يشير إلى أن المخرجات عالية الجودة لا تضمن قابلية قراءة عالية. يستنتج المؤلفون أن تحسين التعليم حول سرطان الجلد الذي يتم إنشاؤه بواسطة الذكاء الاصطناعي يتطلب نهجًا شاملاً يأخذ في الاعتبار كل من اختيار النموذج وموضوع المحتوى المحدد.
مقدمة
تسلط المقدمة الضوء على الدور الحاسم للتعليم الصحي العام في الوقاية من الأمراض وإدارتها، مشددة على ضرورة المعلومات التي تكون دقيقة وقابلة للتنفيذ ومتاحة لمختلف الجماهير. يعد سرطان الجلد مثالاً ملائماً، حيث يؤثر الوعي العام بشكل كبير على الكشف المبكر ونتائج العلاج. يؤكد النص على التحديات المستمرة التي يواجهها المراسلون الصحيون في تحقيق التوازن بين جودة المعلومات—التي تتميز بالوضوح والفائدة والموثوقية—مع الحاجة إلى قابلية قراءة النص، والتي تؤثر على الوصول المعرفي. يتطلب تحقيق مواد تلبي كل من الجودة وقابلية القراءة خبرة كبيرة وعملية تحسين تكرارية.
طرق
توضح قسم “المواد والطرق” تصميم التجربة والإجراءات المستخدمة في الدراسة. يتناول المواد المحددة المستخدمة، بما في ذلك أي مواد كيميائية، معدات، وعينات بيولوجية، لضمان إمكانية تكرار التجارب. تشمل المنهجية البروتوكولات المتبعة لجمع البيانات، بما في ذلك أي تحليلات إحصائية تم تطبيقها لتفسير النتائج.
بالإضافة إلى ذلك، قد يصف القسم إعداد التجربة، بما في ذلك الضوابط والمتغيرات، فضلاً عن التقنيات المستخدمة للقياس والملاحظة. يسمح هذا النهج الشامل بفهم واضح لكيفية إجراء البحث، مما يسهل التحقق واستكشاف النتائج المقدمة في الدراسة بشكل أكبر.
نتائج
يقدم قسم “النتائج” النتائج الرئيسية للدراسة، مسلطًا الضوء على النتائج المهمة المستمدة من الإجراءات التجريبية أو التحليلية المستخدمة. تشير البيانات إلى وجود ارتباط واضح بين المتغيرات المستقلة والتابعة، مع تأكيد التحليلات الإحصائية على قوة هذه العلاقات. من الجدير بالذكر أن النتائج تظهر أن التدخل المطبق يؤدي إلى تحسين قابل للقياس في المقاييس المستهدفة، مع قيمة p أقل من 0.05 تشير إلى الأهمية الإحصائية.
علاوة على ذلك، تكشف النتائج أن ظروفًا أو معايير معينة تؤثر بشكل كبير على فعالية التدخل، مما يشير إلى طرق محتملة لمزيد من البحث. يتم توضيح النتائج من خلال أشكال وجداول متنوعة، والتي توفر تمثيلًا بصريًا لاتجاهات البيانات وتدعم الاستنتاجات المستخلصة. بشكل عام، تؤكد النتائج على أهمية مساهمات الدراسة في الجسم المعرفي القائم في هذا المجال.
مناقشة
تستكشف هذه الدراسة فعالية نماذج اللغة الكبيرة المختلفة باللغة الصينية (LLMs) في إنشاء مواد تعليمية حول سرطان الجلد للقراء العاديين. تشمل النماذج التي تم تقييمها دو باو، ديب سيك، وين شين يي يان، تونغ يي تشيان وين، وGPT-5، مع التركيز على جودة المحتوى وقابلية القراءة. تستخدم الدراسة النسخة الصينية من أداة تقييم مواد التعليم للمرضى (c-PEMAT-P) ومقياس الجودة العالمي (GQS) لتقييم الفهم والجودة العامة، على التوالي. تشير النتائج إلى أنه بينما حقق GPT-5 أعلى درجة في GQS، مما يشير إلى جودة محتوى متفوقة، لم تختلف درجات c-PEMAT بشكل كبير بين النماذج، مما يشير إلى قابلية فهم متقاربة. من الجدير بالذكر أن قابلية القراءة تفاوتت بشكل كبير عبر النماذج وفئات المحتوى، مع عدم تفوق أي نموذج واحد في جميع المقاييس.
تكشف النتائج عن ارتباط ضعيف بين جودة المحتوى وقابلية القراءة، مما يشير إلى أن التحسينات في أحدهما لا تعزز بالضرورة الآخر. يبرز هذا الحاجة إلى استراتيجيات مخصصة في استخدام LLMs للتعليم الصحي، حيث قد تظل المحتويات عالية الجودة معقدة وأقل وصولاً. تؤكد الدراسة على أهمية معالجة كلا البعدين بشكل مستقل لتحسين مواد التعليم للمرضى، خاصة في سياق حيث غالبًا ما تتجاوز الموارد الحالية مستويات القراءة الموصى بها. بشكل عام، بينما تظهر LLMs وعدًا في إنشاء محتوى تعليمي عالي الجودة، فإن الاعتبار الدقيق لكل من الجودة وقابلية القراءة أمر ضروري للتواصل الفعال في الصحة العامة.
DOI: https://doi.org/10.3389/fpubh.2026.1777577
PMID: https://pubmed.ncbi.nlm.nih.gov/41799487
Publication Date: 2026-02-20
Author(s): Yanping Zhang et al.
Primary Topic: Artificial Intelligence in Healthcare and Education
Overview
In this study, the authors investigate the quality and readability of health information generated by five large language models (LLMs)—Doubao, DeepSeek, Wenxin Yiyan, Tongyi Qianwen, and GPT-5—using a standardized set of 20 frequently asked questions about skin cancer in Mandarin Chinese. The evaluation employed the consumer-friendly Patient Education Materials Assessment Tool for Print (c-PEMAT-P) and the Global Quality Scale (GQS) to assess content quality, while readability was measured using seven different indices, including the Automated Readability Index (ARI) and the Flesch Reading Ease Score (FRES).
The findings reveal that while the models exhibited similar levels of understandability and actionability as per c-PEMAT-P, there were significant differences in overall quality, with GPT-5 achieving the highest GQS score. Readability varied considerably across both models and content categories, indicating that no single model excelled across all readability measures. Notably, correlation analyses demonstrated that quality and readability are largely independent, suggesting that high-quality outputs do not guarantee high readability. The authors conclude that enhancing AI-generated skin cancer education necessitates a comprehensive approach that considers both the choice of model and the specific content topic.
Introduction
The introduction highlights the critical role of public health education in disease prevention and management, emphasizing the necessity for information that is accurate, actionable, and accessible to various audiences. Skin cancer serves as a pertinent example, as public awareness significantly impacts early detection and treatment outcomes. The text underscores the ongoing challenges faced by health communicators in balancing the quality of information—characterized by clarity, utility, and reliability—with the need for textual readability, which affects cognitive accessibility. Achieving materials that satisfy both quality and readability requires considerable expertise and a process of iterative refinement.
Methods
The “Materials and Methods” section outlines the experimental design and procedures employed in the study. It details the specific materials used, including any reagents, equipment, and biological samples, ensuring reproducibility of the experiments. The methodology encompasses the protocols followed for data collection, including any statistical analyses applied to interpret the results.
Additionally, the section may describe the experimental setup, including controls and variables, as well as the techniques utilized for measurement and observation. This comprehensive approach allows for a clear understanding of how the research was conducted, facilitating validation and further exploration of the findings presented in the study.
Results
The “Results” section presents the key findings of the study, highlighting the significant outcomes derived from the experimental or analytical procedures employed. The data indicates a clear correlation between the independent and dependent variables, with statistical analyses confirming the robustness of these relationships. Notably, the results demonstrate that the intervention applied leads to a measurable improvement in the target metrics, with a p-value of less than 0.05 indicating statistical significance.
Furthermore, the findings reveal that certain conditions or parameters significantly influence the effectiveness of the intervention, suggesting potential avenues for further research. The results are illustrated through various figures and tables, which provide a visual representation of the data trends and support the conclusions drawn. Overall, the results underscore the importance of the study’s contributions to the existing body of knowledge in the field.
Discussion
This study investigates the effectiveness of various Chinese-language Large Language Models (LLMs) in generating skin cancer educational materials for lay readers. The models evaluated include Doubao, DeepSeek, Wenxin Yiyan, Tongyi Qianwen, and GPT-5, with a focus on their content quality and readability. The research employs the Chinese version of the Patient Education Materials Assessment Tool (c-PEMAT-P) and the Global Quality Scale (GQS) to assess understandability and overall quality, respectively. Results indicate that while GPT-5 achieved the highest GQS score, suggesting superior content quality, the c-PEMAT scores did not significantly differ among the models, indicating comparable understandability. Notably, readability varied significantly across models and content categories, with no single model excelling across all metrics.
The findings reveal a weak correlation between content quality and readability, suggesting that improvements in one do not necessarily enhance the other. This highlights the need for tailored strategies in deploying LLMs for health education, as high-quality content may still be complex and less accessible. The study underscores the importance of addressing both dimensions independently to optimize patient education materials, particularly in a context where existing resources often exceed recommended reading levels. Overall, while LLMs show promise in generating high-quality educational content, careful consideration of both quality and readability is essential for effective public health communication.
