تقييم فائدة ChatGPT في تبسيط تعقيد النصوص لمواد التعليم الخاصة بالمرضى
Assessing the Utility of ChatGPT in Simplifying Text Complexity of Patient Educational Materials

المجلة: Cureus، المجلد: 16، العدد: 3
DOI: https://doi.org/10.7759/cureus.55304
PMID: https://pubmed.ncbi.nlm.nih.gov/38559518
تاريخ النشر: 2024-03-01
المؤلف: Rasika Sudharshan وآخرون
الموضوع الرئيسي: محو الأمية الصحية وإمكانية الوصول إلى المعلومات

نظرة عامة

تدرس الدراسة فعالية ChatGPT النسخة 3.5 في تبسيط المعلومات الصحية من الأكاديمية الأمريكية لطب العيون (AAO) للمرضى، مع التركيز على كل من النصوص الإنجليزية والإسبانية. تم اختيار خمسة عشر مقالًا ومعالجتها من خلال ChatGPT لتقييم تغييرات القابلية للقراءة باستخدام مقاييس مختلفة، بما في ذلك مقياس سهولة القراءة Flesch ومؤشر Gunning Fog للغة الإنجليزية، والعديد من المؤشرات للغة الإسبانية. أشارت النتائج إلى أنه بينما لم تتحسن قابلية قراءة المقالات الإنجليزية بشكل كبير (من مستوى صف دراسي متوسط 8.43 إلى 8.9، p=0.41)، كان هناك تبسيط كبير في المقالات الإسبانية (من 5.3 إلى 4.1، p=0.0001)، ولاحظ بشكل خاص في مقياس Legibilidad-µ (p=0.003).

تشير النتائج إلى أن ChatGPT قد لا يعزز قابلية قراءة المواد الصحية باللغة الإنجليزية ولكنه يظهر وعدًا في تحسين محو الأمية الصحية للمرضى الناطقين بالإسبانية. تدعو الدراسة إلى التوصية باستخدام ChatGPT للأفراد الناطقين بالإسبانية الذين يسعون لفهم مقالات طب العيون من AAO بشكل أفضل. كما تؤكد على الحاجة إلى مزيد من البحث لاستكشاف إمكانيات الدردشة الآلية في تعزيز التواصل عبر مجموعات المرضى المتنوعة ومواضيع طبية مختلفة.

مقدمة

تسلط المقدمة الضوء على الدور الحاسم للإنترنت كمصدر رئيسي للمعلومات الطبية لحوالي 75% من الأفراد في الولايات المتحدة. ومع ذلك، تشير إلى وجود حواجز كبيرة أمام محو الأمية الصحية، بما في ذلك المصطلحات الطبية المعقدة، مستويات القراءة المنخفضة، والقدرة المحدودة على اللغة الإنجليزية، خاصة بين المتحدثين غير الأصليين باللغة الإنجليزية. يمثل ظهور الدردشة الآلية المدعومة بالذكاء الاصطناعي مثل ChatGPT حلاً محتملاً لهذه التحديات من خلال تسهيل الوصول إلى المعلومات الطبية وتعزيز فهم الموارد الموجودة.

تهدف الدراسة إلى التحقيق في فعالية ChatGPT، وبشكل خاص نسخته المجانية 3.5، في تبسيط المفاهيم الطبية للمرضى، وخاصة أولئك الذين لديهم قدرة محدودة على اللغة الإنجليزية. تؤكد على أهمية تحسين محو الأمية الصحية، خاصة في الفئات التي تواجه عوائق تعليمية، كما يتضح من الإحصائيات التي تظهر أن جزءًا كبيرًا من المرضى الناطقين بالإسبانية لديهم مستوى تعليمي منخفض. من خلال دراسة مزايا وتحديات التواصل القائم على الدردشة الآلية في الرعاية الصحية، تسعى الورقة إلى توضيح كيف يمكن أن تحول هذه التكنولوجيا محو الأمية الصحية وتحسن نتائج المرضى.

الطرق

في هذه الدراسة، قام المؤلفون بتقييم فعالية ChatGPT النسخة 3.5 في تعزيز قابلية قراءة المواد التعليمية من الأكاديمية الأمريكية لطب العيون (AAO) باللغتين الإنجليزية والإسبانية. بدأوا التقييم بمطالبة أولية لضمان فهم ChatGPT لمهمة تبسيط النص من أجل تحسين الفهم. تم معالجة خمسة عشر مقالًا تعليميًا تغطي مواضيع عينية مختلفة من خلال ChatGPT باستخدام مطالبات مصممة لاستنباط نسخ مبسطة من النصوص. تم توجيه كل مقال عشر مرات، مما أدى إلى توليد استجابات فريدة للتحليل.

لتحديد قابلية القراءة، استخدم الباحثون صيغ تسجيل القابلية للقراءة المعتمدة والمخصصة لكل لغة. بالنسبة للنصوص الإنجليزية، استخدموا مقياس سهولة القراءة Flesch (FRE) ومؤشر Gunning Fog (GFI)، بينما تم تطبيق Crawford Nivel-de-Grado (CNG) على النصوص الإسبانية. تقيم درجات FRE وGFI عوامل مثل طول الجملة وعدد المقاطع، حيث تشير الدرجات الأعلى إلى سهولة القراءة. كما حسبت الدراسة مستويات الصف باستخدام مستوى الصف Flesch Kincaid (FKGL) للغة الإنجليزية وCNG للغة الإسبانية، حيث تشير القيم الأعلى إلى زيادة الصعوبة. تم إجراء تحليلات إحصائية، بما في ذلك المتوسط والانحراف المعياري، واختبارات t ذات الطرفين، لمقارنة درجات قابلية القراءة لمخرجات ChatGPT مقابل المقالات الأصلية، مع إجراء جميع الحسابات في Excel.

النتائج

تشير النتائج إلى أن متوسط مستوى الصف للمقالات الإنجليزية الأصلية الـ 15 كان $8.43 \pm 1.17$ كما تم تقييمه بواسطة صيغة مستوى الصف Flesch-Kincaid (FKGL). بعد التبسيط باستخدام ChatGPT، زاد مستوى الصف إلى $8.80 \pm 1.88$، مع قيمة p تساوي 0.53، مما يشير إلى عدم وجود تغيير ذو دلالة إحصائية. كان متوسط درجة سهولة القراءة Flesch (FRE) للمقالات الأصلية $61.8 \pm 6.96$، والتي انخفضت إلى $55.97 \pm 12.38$ بعد التبسيط (p=0.12). بالإضافة إلى ذلك، كانت درجات مؤشر Gunning Fog (GFI) $10.81 \pm 1.41$ للمقالات الأصلية و$10.29 \pm 1.56$ للإصدارات المبسطة (p=0.34).

بشكل عام، كشفت التحليلات أنه لم تظهر أي من مقاييس القابلية للقراءة—FKGL، FRE، أو GFI—اختلافات ذات دلالة إحصائية بين المقالات الأصلية والمبسطة. تم تقديم نتائج مقياس القابلية للقراءة التفصيلية لكل مقال في الشكل 1 والجدول 1.

المناقشة

تسلط قسم المناقشة في ورقة البحث الضوء على التأثير المختلف لـ ChatGPT على قابلية قراءة مواد تعليم المرضى باللغة الإسبانية مقابل الإنجليزية. وجدت الدراسة أنه بينما تحسنت قابلية قراءة المقالات الإسبانية بشكل كبير بعد التبسيط (p=0.0001)، لم تظهر المقالات الإنجليزية تغييرًا ذو دلالة إحصائية في القابلية للقراءة، حيث تشير قيم p إلى عدم وجود اختلافات ذات مغزى عبر اختبارات القابلية للقراءة المختلفة. تثير هذه الفجوة تساؤلات حول قدرات ChatGPT، خاصة بالنظر إلى تدريبه الأساسي باللغة الإنجليزية، وتقترح أن الاختلافات الهيكلية في القواعد والنحو بين اللغتين قد تؤثر على فعالية عملية التبسيط.

يشير المؤلفون إلى أن الدراسات السابقة أظهرت قدرة ChatGPT على توليد معلومات دقيقة لتعليم المرضى عبر مجالات طبية متنوعة، ومع ذلك تكشف هذه الدراسة عن قيود في قدراته على تعزيز القابلية للقراءة للنصوص الإنجليزية. تؤكد النتائج على ضرورة إجراء مزيد من البحث لاستكشاف إمكانيات الدردشة الآلية في تحسين محو الأمية الصحية، خاصة بين السكان الناطقين بالإسبانية. يوصي المؤلفون بالاستمرار في تقييم أداء ChatGPT، خاصة مع الإصدارات المستقبلية، ويقترحون أن ملاحظات المرضى وتحليل الخبراء يمكن أن توفر رؤى إضافية حول فعالية المواد التي تنتجها الدردشة الآلية. بشكل عام، بينما تشير النتائج الحالية إلى أن ChatGPT النسخة 3.5 قد يكون مفيدًا للمرضى الناطقين بالإسبانية، فإن مزيدًا من التحقيق مطلوب لتقييم قابليته عبر مواضيع طبية متنوعة وخصائص المرضى المختلفة.

Journal: Cureus, Volume: 16, Issue: 3
DOI: https://doi.org/10.7759/cureus.55304
PMID: https://pubmed.ncbi.nlm.nih.gov/38559518
Publication Date: 2024-03-01
Author(s): Rasika Sudharshan et al.
Primary Topic: Health Literacy and Information Accessibility

Overview

The study investigates the effectiveness of ChatGPT version 3.5 in simplifying healthcare information from the American Academy of Ophthalmology (AAO) for patients, focusing on both English and Spanish texts. Fifteen articles were selected and processed through ChatGPT to assess readability changes using various scales, including the Flesch Reading Ease and Gunning Fog Index for English, and multiple indices for Spanish. The results indicated that while the readability of English articles did not significantly improve (from an average grade level of 8.43 to 8.9, p=0.41), there was a significant simplification in Spanish articles (from 5.3 to 4.1, p=0.0001), particularly noted in the Legibilidad-µ scale (p=0.003).

The findings suggest that ChatGPT may not enhance the readability of English-language healthcare materials but shows promise in improving health literacy for Spanish-speaking patients. The study advocates for the recommendation of ChatGPT for Spanish-speaking individuals seeking to understand AAO ophthalmology articles better. It also emphasizes the need for further research to explore the potential of chatbots in enhancing communication across diverse patient populations and medical topics.

Introduction

The introduction highlights the critical role of the Internet as a primary source of medical information for approximately 75% of individuals in the US. However, it notes significant barriers to health literacy, including complex medical jargon, lower reading levels, and limited English proficiency, particularly among non-native English speakers. The emergence of AI chatbots like ChatGPT presents a potential solution to these challenges by facilitating easier access to medical information and enhancing comprehension of existing resources.

The study aims to investigate the effectiveness of ChatGPT, specifically its free version 3.5, in simplifying medical concepts for patients, particularly those with limited English proficiency. It underscores the importance of improving health literacy, especially in populations that face educational disadvantages, as evidenced by statistics showing that a significant portion of Spanish-speaking patients have low educational attainment. By examining the advantages and challenges of chatbot-based communication in healthcare, the paper seeks to illuminate how such technology can transform health literacy and improve patient outcomes.

Methods

In this study, the authors evaluated the effectiveness of ChatGPT version 3.5 in enhancing the readability of educational materials from the American Academy of Ophthalmology (AAO) in both English and Spanish. They initiated the assessment with a preliminary prompt to ensure ChatGPT understood the task of simplifying text for better comprehension. A total of fifteen educational articles covering various ophthalmological topics were processed through ChatGPT using prompts designed to elicit simplified versions of the texts. Each article was prompted ten times, generating unique responses for analysis.

To quantify readability, the researchers employed established readability scoring formulas tailored for each language. For English texts, they utilized the Flesch Reading Ease (FRE) and Gunning Fog Index (GFI), while for Spanish texts, the Crawford Nivel-de-Grado (CNG) was applied. The FRE and GFI scores assess factors such as sentence length and syllable count, with higher scores indicating easier readability. The study also calculated grade levels using the Flesch Kincaid Grade Level (FKGL) for English and CNG for Spanish, where higher values denote increased difficulty. Statistical analyses, including mean, standard deviation, and two-tailed t-tests, were conducted to compare the readability scores of ChatGPT outputs against the original articles, with all calculations performed in Excel.

Results

The results indicate that the average grade level of the 15 original English articles was $8.43 \pm 1.17$ as assessed by the Flesch-Kincaid Grade Level (FKGL) formula. After simplification using ChatGPT, the grade level increased to $8.80 \pm 1.88$, with a p-value of 0.53, suggesting no statistically significant change. The Flesch Reading Ease (FRE) score for the original articles averaged $61.8 \pm 6.96$, which decreased to $55.97 \pm 12.38$ post-simplification (p=0.12). Additionally, the Gunning Fog Index (GFI) scores were $10.81 \pm 1.41$ for the original articles and $10.29 \pm 1.56$ for the simplified versions (p=0.34).

Overall, the analysis revealed that none of the readability measures—FKGL, FRE, or GFI—demonstrated statistically significant differences between the original and simplified articles. Detailed readability scale results for each article are presented in Figure 1 and Table 1.

Discussion

The discussion section of the research paper highlights the differential impact of ChatGPT on the readability of patient education materials in Spanish versus English. The study found that while the readability of Spanish articles significantly improved post-simplification (p=0.0001), the English articles did not show a statistically significant change in readability, with p-values indicating no meaningful differences across various readability tests. This discrepancy raises questions about the capabilities of ChatGPT, particularly given its primary training in English, and suggests that the structural differences in grammar and syntax between the two languages may influence the effectiveness of the simplification process.

The authors note that previous studies have demonstrated ChatGPT’s ability to generate accurate patient education information across various medical fields, yet this study reveals limitations in its readability enhancement capabilities for English texts. The findings underscore the necessity for further research to explore the potential of chatbots in improving health literacy, particularly among Spanish-speaking populations. The authors recommend continued evaluation of ChatGPT’s performance, especially with future versions, and suggest that patient feedback and expert analysis could provide additional insights into the effectiveness of chatbot-generated materials. Overall, while the current results indicate that ChatGPT version 3.5 may be beneficial for Spanish-speaking patients, further investigation is warranted to assess its applicability across diverse medical topics and patient demographics.