تقييم موثوقية وقابلية قراءة استجابات ChatGPT-4 بشأن قصور الغدة الدرقية أثناء الحمل
Evaluation of the reliability and readability of ChatGPT-4 responses regarding hypothyroidism during pregnancy

المجلة: Scientific Reports، المجلد: 14، العدد: 1
DOI: https://doi.org/10.1038/s41598-023-50884-w
PMID: https://pubmed.ncbi.nlm.nih.gov/38167988
تاريخ النشر: 2024-01-02
المؤلف: Çağatay Emir Önder وآخرون
الموضوع الرئيسي: الذكاء الاصطناعي في الرعاية الصحية والتعليم

نظرة عامة

تقيّم هذه الدراسة موثوقية وقراءة الاستجابات التي تم إنشاؤها بواسطة ChatGPT-4 بشأن قصور الغدة الدرقية أثناء الحمل، وهي حالة معروفة بتأثيرها السلبي على صحة الأم والجنين. تم طرح ما مجموعه 19 سؤالًا متوافقًا مع إرشادات الجمعية الأمريكية للغدة الدرقية على ChatGPT-4، وتم تقييم الاستجابات من حيث الجودة باستخدام مقياس الجودة العالمي (GQS) وأدوات DISCERN المعدلة. أشارت النتائج إلى عدم وجود معلومات مضللة، مع متوسط درجة DISCERN المعدلة 30.26 ± 3.14 ودرجة GQS الوسيطة 4 (النطاق 2-4). أظهرت معظم الاستجابات موثوقية متوسطة (78.9%) إلى جيدة (21.1%). ومع ذلك، كشفت تقييمات القراءة عن درجة Flesch Reading Ease (FRE) الوسيطة 32.20، مما يشير إلى مستوى قراءة يتطلب عادةً تعليمًا جامعيًا.

في الختام، بينما يوفر ChatGPT-4 معلومات موثوقة حول قصور الغدة الدرقية أثناء الحمل، فإن قابلية القراءة تمثل عائقًا لفهم الجمهور العام. تقترح الدراسة أنه، على الرغم من إمكانيات النموذج كمصدر معلومات مساعد للتواصل بين المرضى والأطباء، فإن تحسينات في كل من الموثوقية وقابلية القراءة ضرورية. بالإضافة إلى ذلك، يُوصى بمزيد من البحث لاستكشاف أداء ChatGPT-4 عبر لغات مختلفة ومع مجموعة أوسع من الأسئلة، فضلاً عن تطوير نماذج محددة المجال لتعزيز فائدته.

طرق

يستعرض قسم “المواد والطرق” في ورقة البحث التصميم التجريبي والإجراءات المستخدمة للتحقيق في سؤال البحث. يوضح المواد المحددة المستخدمة، بما في ذلك أي مواد كيميائية أو أدوات أو تقنيات، ويصف البروتوكولات المتبعة لضمان إمكانية تكرار النتائج وموثوقيتها. قد يتضمن القسم أيضًا معلومات حول أحجام العينات، وظروف التحكم، والتحليلات الإحصائية التي تم إجراؤها لتفسير البيانات.

بالإضافة إلى ذلك، فإن الطرق المستخدمة ضرورية للتحقق من النتائج وضمان أن النتائج قوية. من المحتمل أن يبرز القسم المنطق وراء المنهجيات المختارة، موضحًا كيف تتماشى مع أهداف الدراسة. بشكل عام، يعد هذا الجزء من الورقة عنصرًا أساسيًا، يمكّن الباحثين الآخرين من تكرار الدراسة والبناء على نتائجها.

نتائج

يقدم قسم “النتائج” في ورقة البحث النتائج الرئيسية المستمدة من التجارب أو التحليلات التي تم إجراؤها. يبرز الاتجاهات البيانية المهمة، والنتائج الإحصائية، وأي علاقات ملحوظة بين المتغيرات. عادةً ما تكون النتائج مصحوبة بأشكال وجداول ذات صلة توضح البيانات بصريًا، مما يسهل تفسير النتائج.

قد يناقش القسم أيضًا تداعيات النتائج فيما يتعلق بالفرضيات المطروحة في الدراسة، مشيرًا إلى ما إذا كانت البيانات تدعم أو تنفي الافتراضات الأولية. يتم ملاحظة أي نتائج غير متوقعة أو شذوذ، مما يوفر نظرة شاملة على مساهمات البحث في هذا المجال. بشكل عام، يعد هذا القسم عنصرًا حاسمًا في فهم تداعيات نتائج الدراسة في سياق الأدبيات الموجودة.

مناقشة

في هذه الدراسة، تم تقييم موثوقية وقراءة الاستجابات التي تم إنشاؤها بواسطة ChatGPT-4 بشأن قصور الغدة الدرقية أثناء الحمل من خلال سلسلة من الأسئلة المفتوحة وسيناريوهات المرضى. أشارت النتائج إلى أن ChatGPT-4 قدم استجابات بموثوقية متوسطة إلى جيدة وجودة عالية، مع متوسط درجة توافقية 30.26 على مقياس DISCERN المعدل (mDISCERN) ودرجة جودة عالمية (GQS) وسيطة 4. ومن الجدير بالذكر أنه لم يتم العثور على معلومات مضللة في أي من الاستجابات. ومع ذلك، تم تقييم قابلية قراءة النص على أنها صعبة، مما يتطلب مستوى تعليمي جامعي لفهمها، مما قد يحد من إمكانية الوصول للجمهور العام.

تؤكد الدراسة على أهمية المعلومات الدقيقة بشأن قصور الغدة الدرقية أثناء الحمل، نظرًا للمخاطر المرتبطة بالنتائج السلبية مثل الولادة المبكرة وانخفاض وزن الولادة. بينما أظهر ChatGPT-4 إمكانيات كمصدر موثوق للمعلومات، فإن تعقيد لغته يتطلب تحسينات في قابلية القراءة لتعزيز فائدته للمرضى. يقترح المؤلفون أن البحث المستقبلي يجب أن يستكشف أداء ChatGPT-4 في لغات وسياقات متنوعة، فضلاً عن تطوير نماذج محددة المجال لخدمة كل من الأطباء والمرضى بشكل أفضل.

Journal: Scientific Reports, Volume: 14, Issue: 1
DOI: https://doi.org/10.1038/s41598-023-50884-w
PMID: https://pubmed.ncbi.nlm.nih.gov/38167988
Publication Date: 2024-01-02
Author(s): Çağatay Emir Önder et al.
Primary Topic: Artificial Intelligence in Healthcare and Education

Overview

This study evaluates the reliability and readability of responses generated by ChatGPT-4 regarding hypothyroidism in pregnancy, a condition known to adversely affect both maternal and fetal health. A total of 19 questions aligned with the American Thyroid Association guidelines were posed to ChatGPT-4, and the responses were assessed for quality using the global quality scale (GQS) and modified DISCERN tools. The findings indicated that no misleading information was present, with a mean modified DISCERN score of 30.26 ± 3.14 and a median GQS score of 4 (range 2-4). Most responses exhibited moderate (78.9%) to good (21.1%) reliability. However, readability assessments revealed a median Flesch Reading Ease (FRE) score of 32.20, indicating a reading level that typically requires a college education.

In conclusion, while ChatGPT-4 provides reliable information on hypothyroidism in pregnancy, its readability poses a barrier for general public understanding. The study suggests that, despite the model’s potential as an auxiliary information source for patient-clinician communication, improvements in both reliability and readability are necessary. Additionally, further research is recommended to explore ChatGPT-4’s performance across different languages and with a broader range of questions, as well as the development of domain-specific models to enhance its utility.

Methods

The “Materials and Methods” section of the research paper outlines the experimental design and procedures employed to investigate the research question. It details the specific materials used, including any reagents, instruments, or technologies, and describes the protocols followed to ensure reproducibility and reliability of the results. The section may also include information on sample sizes, control conditions, and statistical analyses performed to interpret the data.

Additionally, the methods employed are crucial for validating the findings and ensuring that the results are robust. The section likely emphasizes the rationale behind the chosen methodologies, highlighting how they align with the objectives of the study. Overall, this part of the paper serves as a foundational element, enabling other researchers to replicate the study and build upon its findings.

Results

The “Results” section of the research paper presents key findings derived from the conducted experiments or analyses. It highlights significant data trends, statistical outcomes, and any observed relationships between variables. The results are typically accompanied by relevant figures and tables that illustrate the data visually, allowing for easier interpretation of the findings.

The section may also discuss the implications of the results in relation to the hypotheses posed in the study, indicating whether the data supports or refutes the initial assumptions. Any unexpected outcomes or anomalies are noted, providing a comprehensive overview of the research’s contributions to the field. Overall, this section serves as a critical component in understanding the implications of the study’s findings within the broader context of existing literature.

Discussion

In this study, the reliability and readability of responses generated by ChatGPT-4 regarding hypothyroidism during pregnancy were evaluated through a series of open-ended questions and patient scenarios. The findings indicated that ChatGPT-4 provided responses with moderate to good reliability and high quality, with a consensus mean score of 30.26 on the modified DISCERN (mDISCERN) scale and a median Global Quality Score (GQS) of 4. Notably, no misleading information was found in any of the responses. However, the readability of the text was assessed as difficult, requiring a college-level education for comprehension, which may limit accessibility for the general public.

The study underscores the importance of accurate information regarding hypothyroidism in pregnancy, given the associated risks of adverse outcomes such as preterm birth and low birth weight. While ChatGPT-4 has shown potential as a reliable source of information, the complexity of its language necessitates improvements in readability to enhance its utility for patients. The authors suggest that future research should explore the performance of ChatGPT-4 in various languages and contexts, as well as the development of domain-specific models to better serve both clinicians and patients.