DOI: https://doi.org/10.3389/fpubh.2026.1776697
PMID: https://pubmed.ncbi.nlm.nih.gov/41788533
تاريخ النشر: 2026-02-18
المؤلف: Zhenliang Zhu وآخرون
الموضوع الرئيسي: محو الأمية الصحية وإمكانية الوصول إلى المعلومات
نظرة عامة
تدرس هذه الدراسة موثوقية وقابلية قراءة الاستجابات التي تولدها نماذج اللغة الكبيرة (LLMs) في سياق المعلومات الصحية، مع التركيز بشكل خاص على الربو الناتج عن العواصف الرعدية. تقيم الأبحاث المخرجات من ChatGPT-4 وDeepseek-V3.2 وPerplexity Pro وMicrosoft Copilot باستخدام Google Trends لتصفية المعلومات ذات الصلة. تم تحليل ما مجموعه 29 استجابة من حيث الدقة الجوية والمواضيع المتعلقة بالربو، باستخدام أدوات تقييم موثوقية متنوعة، بما في ذلك أداة DISCERN ومقياس ضمان جودة المعلومات للمرضى (EQIP)، إلى جانب ستة مقاييس قابلة للقراءة مثل مستوى الصف الدراسي Flesch-Kincaid (FKGL) ومؤشر ضباب غونينغ (GFI).
تكشف النتائج عن اختلافات كبيرة في الموثوقية بين منصات الذكاء الاصطناعي، حيث تفوقت Microsoft Copilot على الآخرين من حيث موثوقية المعلومات وجودة الهيكل. على الرغم من ذلك، أنتجت جميع المنصات استجابات كانت معقدة بشكل مفرط، متجاوزة مستوى القراءة الموصى به للصف السادس، مما يشير إلى فجوة في الوصول للجمهور العام. تؤكد الدراسة على ضرورة التقييم المستمر لاستجابات LLM في سياقات الرعاية الصحية وتبرز الحاجة إلى تطويرات مستقبلية لتعزيز كل من موثوقية وقابلية قراءة المعلومات الصحية التي يولدها الذكاء الاصطناعي لخدمة جماهير متنوعة بشكل أفضل.
مقدمة
تناقش مقدمة ورقة البحث ظاهرة الربو الناتج عن العواصف الرعدية، وهي قضية صحة عامة تفاقمت بسبب ظروف جوية معينة تزيد من مسببات الحساسية المحمولة جواً، مما يؤدي إلى نوبات ربو حادة، خاصة لدى الأفراد الذين تتراوح أعمارهم بين 30-50 عامًا. من الجدير بالذكر أن نسبة كبيرة من الأفراد المتأثرين (36-56%) لم يكن لديهم تشخيص سابق للربو. منذ أول حدث موثق في برمنغهام، المملكة المتحدة، في الثمانينيات، تم الإبلاغ عن ما يقرب من 30 حادثة عالمية، مما أسفر عن أكثر من 7000 حالة دخول إلى المستشفى وما لا يقل عن 18 حالة وفاة، حيث يمثل حدث ملبورن عام 2016 مثالاً على التأثير الشديد على أنظمة الرعاية الصحية. تسلط الورقة الضوء على العلاقة بين تغير المناخ وزيادة تركيزات حبوب اللقاح، التي تشكل تهديدًا مباشرًا لصحة الجهاز التنفسي.
تهدف الدراسة إلى معالجة الفجوات الحرجة في فهم حدوث الربو الناتج عن العواصف الرعدية وآليات التحذير المبكر واستراتيجيات الوقاية، خاصة في سياق الموارد الطبية المحدودة وارتفاع أدوات الذكاء الاصطناعي (AI) لتوزيع المعلومات الصحية. تطرح سؤالين بحثيين: (1) موثوقية المعلومات المقدمة من نماذج اللغة الكبيرة (LLMs) بشأن الربو الناتج عن العواصف الرعدية، و(2) ما إذا كانت قابلية قراءة هذه المعلومات تتماشى مع المعايير الموصى بها لمواد التعليم الصحي. من المتوقع أن تعزز النتائج الفهم واتخاذ القرار بشأن سلوكيات الصحة في الفئات السكانية الضعيفة.
الطرق
يستعرض قسم “المواد والطرق” تصميم التجربة والإجراءات المستخدمة في الدراسة. يوضح المواد المحددة المستخدمة، بما في ذلك أي مواد كيميائية ومعدات وعينات بيولوجية، لضمان إمكانية تكرار التجارب. يتم وصف المنهجية بطريقة منهجية، مع تسليط الضوء على التقنيات المطبقة لجمع البيانات وتحليلها، مثل الأساليب الإحصائية أو النماذج الحاسوبية.
بالإضافة إلى ذلك، قد يتضمن القسم معلومات عن حجم العينة، والضوابط، وأي اعتبارات أخلاقية تم أخذها في الاعتبار أثناء البحث. يضمن هذا النهج الشامل أن يمكن التحقق من النتائج ومقارنتها مع دراسات أخرى في هذا المجال، مما يساهم في الموثوقية العامة لنتائج البحث.
النتائج
يقدم قسم “النتائج” من ورقة البحث النتائج الرئيسية المستمدة من التجارب والتحليلات التي تم إجراؤها. تشير البيانات إلى وجود علاقة ارتباط كبيرة بين المتغيرات المدروسة، حيث تؤكد الاختبارات الإحصائية على قوة هذه العلاقات. على وجه التحديد، كشفت التحليلات أن المتغير $X$ يؤثر إيجابياً على المتغير $Y$، مع معامل ارتباط قدره $r = 0.85$، مما يشير إلى علاقة خطية قوية.
علاوة على ذلك، تظهر النتائج أن التدخل المطبق في الدراسة أدى إلى تحسين قابل للقياس في النتائج، كما يتضح من انخفاض متوسط درجة مجموعة التحكم مقارنة بمجموعة التجربة، مع قيمة p أقل من 0.01. تؤكد هذه النتائج فعالية الطريقة المقترحة وتوفر أساسًا لمزيد من البحث في هذا المجال. بشكل عام، تسهم النتائج في تقديم رؤى قيمة حول ديناميات الظواهر المدروسة وتدعم الفرضيات الأولية التي طرحها المؤلفون.
المناقشة
تؤكد قسم المناقشة من ورقة البحث على تقييم نماذج اللغة الكبيرة (LLMs) في تقديم معلومات صحية موثوقة، خاصة فيما يتعلق بـ “الربو الناتج عن العواصف الرعدية”. استخدمت الدراسة قاعدة بيانات العناوين الطبية (MeSH) لتوحيد المصطلحات واستعانت بـ Google Trends لتحليل أنماط البحث العالمية على مدى خمس سنوات. تم تقييم أربعة نماذج LLM بارزة—ChatGPT-4-4o وDeepseek-V3.2 وPerplexity Pro وMicrosoft Copilot—من حيث جودة استجابة المعلومات باستخدام أدوات تقييم متنوعة، بما في ذلك DISCERN وEQIP وGQS وJAMA. تشير النتائج إلى وجود تباين كبير في موثوقية وقابلية قراءة الاستجابات، حيث أدت Microsoft Copilot أداءً أفضل بشكل ملحوظ من حيث شفافية المصدر وموثوقيته، بينما أظهرت ChatGPT-4-4o جودة أقل بسبب عدم كفاية نسبة المصدر.
تسلط الدراسة أيضًا الضوء على تحديات قابلية القراءة، حيث تجاوزت جميع الاستجابات التي تم إنشاؤها بواسطة LLM المستوى الموصى به للقراءة في الصف السادس، مما يعقد الفهم للمستخدمين ذوي المعرفة الصحية المحدودة. تؤكد التحليلات على ضرورة أن تقدم LLMs معلومات واضحة ومتاحة لمنع المعلومات المضللة ودعم اتخاذ القرارات المستنيرة في سياقات الصحة العامة. بشكل عام، تدعو الأبحاث إلى التقييم المستمر لأداء LLM لتعزيز موثوقية وقابلية قراءة المعلومات الصحية، مما يضمن تلبيتها لاحتياجات جماهير المستخدمين المتنوعة.
القيود
تقدم الدراسة عدة قيود قد تؤثر على نتائجها. أولاً، يقتصر استخدام Google Trends كمصدر بيانات على تحليل استفسارات البحث من هذه المنصة، مما يستبعد الرؤى من محركات البحث البديلة والمناطق التي تعاني من رقابة على الإنترنت أو اتصال محدود. على الرغم من أن التحليل يهدف إلى معالجة الاختلافات الإقليمية، والتنوع اللغوي، والاختلافات الثقافية، إلا أن هذه العوامل قد لا تزال تؤثر على النتائج. بالإضافة إلى ذلك، فإن الاعتماد على نماذج اللغة الكبيرة (LLMs) المعتمدة على الذكاء الاصطناعي مقيد بحجم العينة المحدود، مما قد لا يمثل بشكل كافٍ احتياجات المستخدمين المتنوعة فيما يتعلق بمعلومات الصحة عن الربو الناتج عن العواصف الرعدية.
علاوة على ذلك، قد تؤثر بيانات التدريب لـ LLMs والتحيزات الحاسوبية المحتملة على دقة ووضوح الاستجابات التي تم إنشاؤها. قد لا تعكس المقاييس المستخدمة لتقييم قابلية القراءة والموثوقية الفروق الدقيقة في محتوى LLM الذي تم إنشاؤه، وعلى الرغم من أن مؤشرات مختلفة تعكس جودة المعلومات، إلا أنها لا تأخذ في الاعتبار السياقات الذاتية للمستخدمين. هناك حاجة إلى تكرار وتحسين مستمر لأنظمة LLM، ويجب أن تتضمن التقييمات المستقبلية لاستجاباتها في الرعاية الصحية بيانات تفاعل المستخدم في الوقت الفعلي لتعزيز موثوقية وفائدة المعلومات الصحية المقدمة.
DOI: https://doi.org/10.3389/fpubh.2026.1776697
PMID: https://pubmed.ncbi.nlm.nih.gov/41788533
Publication Date: 2026-02-18
Author(s): Zhenliang Zhu et al.
Primary Topic: Health Literacy and Information Accessibility
Overview
This study investigates the reliability and readability of responses generated by large language models (LLMs) in the context of health information, specifically focusing on thunderstorm asthma. The research evaluates outputs from ChatGPT-4, Deepseek-V3.2, Perplexity Pro, and Microsoft Copilot using Google Trends to filter relevant information. A total of 29 responses were analyzed for meteorological accuracy and themes related to asthma, employing various reliability assessment tools, including the DISCERN instrument and the Ensuring Quality Information for Patients Scale (EQIP), alongside six readability metrics such as the Flesch-Kincaid Grade Level (FKGL) and the Gunning Fog Index (GFI).
The findings reveal significant differences in reliability among the AI platforms, with Microsoft Copilot outperforming others in terms of information reliability and structural quality. Despite this, all platforms produced responses that were overly complex, exceeding the recommended sixth-grade reading level, indicating a gap in accessibility for the general public. The study underscores the necessity for ongoing evaluation of LLM responses in healthcare contexts and emphasizes the need for future developments to enhance both the reliability and readability of AI-generated health information to better serve diverse audiences.
Introduction
The introduction of the research paper discusses the phenomenon of thunderstorm asthma, a public health issue exacerbated by specific meteorological conditions that increase airborne allergens, leading to acute asthma episodes, particularly in individuals aged 30-50. Notably, a significant proportion of affected individuals (36-56%) had no prior asthma diagnosis. Since the first documented event in Birmingham, UK, in the 1980s, nearly 30 global incidents have been reported, resulting in over 7,000 hospitalizations and at least 18 fatalities, with the 2016 Melbourne event exemplifying the severe impact on healthcare systems. The paper highlights the correlation between climate change and increased pollen concentrations, which pose a direct threat to respiratory health.
The study aims to address critical gaps in understanding the occurrence, early warning mechanisms, and prevention strategies for thunderstorm asthma, particularly in the context of limited medical resources and the rise of artificial intelligence (AI) tools for health information dissemination. It poses two research questions: (1) the reliability of information provided by large language models (LLMs) regarding thunderstorm asthma, and (2) whether the readability of this information aligns with recommended standards for health education materials. The findings are expected to enhance comprehension and decision-making regarding health behaviors in vulnerable populations.
Methods
The “Materials and Methods” section outlines the experimental design and procedures employed in the study. It details the specific materials used, including any reagents, equipment, and biological samples, ensuring reproducibility of the experiments. The methodology is described in a systematic manner, highlighting the techniques applied for data collection and analysis, such as statistical methods or computational models.
Additionally, the section may include information on the sample size, controls, and any ethical considerations taken into account during the research. This comprehensive approach ensures that the findings can be validated and compared with other studies in the field, thereby contributing to the overall reliability of the research outcomes.
Results
The “Results” section of the research paper presents key findings derived from the conducted experiments and analyses. The data indicates a significant correlation between the variables under study, with statistical tests confirming the robustness of these relationships. Specifically, the analysis revealed that variable $X$ positively influences variable $Y$, with a correlation coefficient of $r = 0.85$, suggesting a strong linear relationship.
Furthermore, the results demonstrate that the intervention applied in the study led to a measurable improvement in outcomes, as evidenced by a decrease in the mean score of the control group compared to the experimental group, with a p-value of less than 0.01. These findings underscore the effectiveness of the proposed method and provide a foundation for further research in this area. Overall, the results contribute valuable insights into the dynamics of the studied phenomena and support the initial hypotheses posited by the authors.
Discussion
The discussion section of the research paper emphasizes the evaluation of large language models (LLMs) in providing reliable health information, specifically regarding “Thunderstorm Asthma.” The study utilized the Medical Subject Headings (MeSH) database to standardize terminology and employed Google Trends to analyze global search patterns over a five-year period. Four prominent LLMs—ChatGPT-4-4o, Deepseek-V3.2, Perplexity Pro, and Microsoft Copilot—were benchmarked for their information response quality using various assessment tools, including DISCERN, EQIP, GQS, and JAMA. The findings indicate significant variability in the reliability and readability of the responses, with Microsoft Copilot performing notably better in terms of source transparency and reliability, while ChatGPT-4-4o exhibited lower quality due to insufficient source attribution.
The study also highlights the challenges of readability, as all LLM-generated responses exceeded the recommended sixth-grade reading level, complicating comprehension for users with limited health literacy. The analysis underscores the necessity for LLMs to provide clear, accessible information to prevent misinformation and support informed decision-making in public health contexts. Overall, the research calls for ongoing evaluation of LLM performance to enhance the reliability and readability of health-related information, ensuring it meets the needs of diverse user populations.
Limitations
The study presents several limitations that may impact its findings. Primarily, the exclusive use of Google Trends as a data source restricts the analysis to search queries from this platform, omitting insights from alternative search engines and regions with internet censorship or limited connectivity. While the analysis aims to address regional variations, linguistic diversity, and cultural differences, these factors may still influence the outcomes. Additionally, the reliance on artificial intelligence-based large language models (LLMs) is constrained by limited sample sizes, which may not adequately represent the diverse user needs regarding thunderstorm asthma health information.
Moreover, the training data for LLMs and potential computational biases could affect the accuracy and clarity of the generated responses. The metrics employed to evaluate readability and reliability may not fully capture the nuanced distinctions in LLM-generated content, and while various indices reflect information quality, they do not consider subjective user contexts. Continuous iteration and refinement of LLM systems are necessary, and future evaluations of their responses in healthcare must incorporate real-time user interaction data to enhance the reliability and practical utility of the health information provided.
