DOI: https://doi.org/10.1038/s41598-024-83575-1
PMID: https://pubmed.ncbi.nlm.nih.gov/39799203
تاريخ النشر: 2025-01-11
المؤلف: Meryem Şahin Özdemir وآخرون
الموضوع الرئيسي: تغطية اللقاحات والتردد
نظرة عامة
تعتبر هذه الدراسة الأولى التي تقيم كفاية وموثوقية روبوتات الدردشة ChatGPT وGemini في تقديم معلومات حول التهاب الكبد الفيروسي. تم تصنيف إجمالي 176 سؤالًا إلى ثلاث مجموعات: أسئلة وأجوبة عامة من مركز السيطرة على الأمراض والوقاية منها (CDC)، وتوصيات قوية من الإرشادات الدولية، والأسئلة الشائعة من وسائل التواصل الاجتماعي. تم تقييم ردود كلا الروبوتين من قبل متخصصي الأمراض المعدية باستخدام مقياس تقييم، وتم تقييم موثوقية المقياس باستخدام معامل كوهين. أظهرت النتائج أن كل من ChatGPT وGemini كان لهما درجات متوسطة مشابهة (3.55 ± 0.83 مقابل 3.57 ± 0.89، p = 0.260) ونسب إجابات صحيحة تمامًا (71.0% مقابل 78.4%، p = 0.111). أظهرت تحليلات المجموعات الفرعية أداءً قابلاً للمقارنة عبر جميع فئات الأسئلة، على الرغم من أن الدقة بالنسبة للأسئلة المتعلقة بالإرشادات كانت أقل بشكل ملحوظ.
وجدت الدراسة علاقة إيجابية معتدلة بين الإجابات المقدمة من ChatGPT وGemini (r = 0.633، p < 0.001)، مع معدلات تكرار تبلغ 91.3% لـ ChatGPT و92% لـ Gemini (p = 0.710). أشارت قيم كوهين إلى اتفاق كبير بين المقيمين لكل من الروبوتين (κ = 0.720 لـ ChatGPT وκ = 0.704 لـ Gemini). بينما عالج كلا الروبوتين أسئلة CDC ووسائل التواصل الاجتماعي بشكل فعال، كان أداؤهما على أسئلة الإرشادات غير كافٍ. تشير النتائج إلى أن ChatGPT وGemini يمكن أن يكونا موارد قيمة لمعلومات الصحة العامة حول التهاب الكبد الفيروسي، ولكن هناك حاجة إلى تحسينات لتعزيز موثوقيتهما للمهنيين الصحيين.
الطرق
هدفت هذه الدراسة، التي أجريت من 1 مارس إلى 30 مارس 2024، إلى تقييم جودة الردود التي تم إنشاؤها بواسطة ChatGPT وGoogle Gemini بشأن التهاب الكبد الفيروسي. شملت البحث ثلاث مجموعات من الأسئلة: (1) “أسئلة وأجوبة (Q&As) للجمهور” من مراكز السيطرة على الأمراض والوقاية منها (CDC)، (2) توصيات قائمة على الأدلة من الإرشادات الدولية حول التهاب الكبد الفيروسي، و(3) أسئلة شائعة مأخوذة من منصات وسائل التواصل الاجتماعي ومواقع جمعيات الرعاية الصحية. تم تضمين إجمالي 176 سؤالًا، مع 34.7% من CDC، و42.6% من الإرشادات، و22.7% من وسائل التواصل الاجتماعي. تم تطبيق معايير صارمة لاستبعاد الأسئلة غير ذات الصلة أو غير الواضحة.
تم تقييم ردود ChatGPT 3.5 وGoogle Gemini 1.5 Flash من قبل اثنين من متخصصي الأمراض المعدية والميكروبيولوجيا السريرية، وتم حل التباينات بواسطة خبير ثالث. تم تقييم الإجابات باستخدام نظام تقييم موحد على مقياس من 1 (خاطئ تمامًا) إلى 4 (صحيح تمامًا وكافٍ). لتقييم القابلية للتكرار، تم طرح كل سؤال على جهازين مختلفين في أوقات منفصلة، وتصنيف الردود على أنها متسقة أو غير متسقة. لم تتطلب الدراسة موافقة لجنة الأخلاقيات حيث لم تتضمن بيانات المرضى.
النتائج
قيمت الدراسة أداء روبوتي الدردشة AI، ChatGPT وGemini، في الإجابة على أسئلة تتعلق بأنواع مختلفة من التهاب الكبد. تم تحليل إجمالي 176 سؤالًا، مع توزيع عبر التهاب الكبد A (18.8%)، B (29.5%)، C (21.0%)، D (17.0%)، وE (13.7%). كان هناك اتفاق كبير بين المقيمين لكل من الروبوتين، مع قيم كوهين تبلغ κ = 0.720 لـ ChatGPT وκ = 0.704 لـ Gemini. حقق ChatGPT معدل إجابات صحيحة بلغ 71.0%، بينما حقق Gemini أداءً أفضل قليلاً بنسبة 78.4%. من الجدير بالذكر أن كلا الروبوتين أظهرا دقة أعلى في الرد على أسئلة CDC مقارنة بأسئلة الإرشادات، حيث سجل ChatGPT 90.1% مقابل 49.4% وGemini سجل 91.9% مقابل 61.4%.
أظهر التحليل الإضافي أن دقة الردود تختلف حسب نوع التهاب الكبد، مع أعلى معدلات إجابات صحيحة لالتهاب الكبد E (91.6%) وأدنى معدل لالتهاب الكبد D (60.0%) لـ ChatGPT. بالنسبة لـ Gemini، كانت المعدلات مرتفعة أيضًا لالتهاب الكبد E (91.6%) وتفاوتت لأنواع أخرى. كانت الدرجات المتوسطة للردود قابلة للمقارنة، حيث سجل ChatGPT 3.55 ± 0.83 وGemini 3.57 ± 0.89، مما يشير إلى عدم وجود فرق كبير (p = 0.260). بالإضافة إلى ذلك، كانت القابلية للتكرار للإجابات عالية لكل من الروبوتين، حيث كانت 91.3% لـ ChatGPT و92% لـ Gemini، مما يشير إلى أداء متسق عبر فئات الأسئلة المختلفة. بشكل عام، بينما قدم كلا الروبوتين أداءً جيدًا، أظهر Gemini معدل دقة أعلى بشكل طفيف.
المناقشة
في هذه الدراسة، تم تقييم موثوقية ومعرفة روبوتات الدردشة ChatGPT وGemini بشأن التهاب الكبد الفيروسي باستخدام التحليل الإحصائي عبر SPSS الإصدار 25.0. تم التعبير عن المتغيرات المستمرة كمتوسط ± انحراف معياري، بينما تم تقديم المتغيرات الفئوية كأعداد ونسب. وجدت الدراسة أن كلا الروبوتين قدما أداءً جيدًا في الإجابة على أسئلة من CDC (90.1% لـ ChatGPT و91.9% لـ Gemini) ومنصات وسائل التواصل الاجتماعي (82.5% لـ ChatGPT و90.0% لـ Gemini). ومع ذلك، كان أداؤهما على الأسئلة المستندة إلى الإرشادات أقل بشكل ملحوظ، مع معدلات إجابات صحيحة بلغت 49.4% لـ ChatGPT و61.4% لـ Gemini. على الرغم من معدلات القابلية للتكرار العالية (91.3% لـ ChatGPT و92% لـ Gemini)، واجه كلا الروبوتين صعوبة في أسئلة توصيات العلاج، حيث حققا دقة بلغت 42.8% فقط.
تشير النتائج إلى أنه بينما يمكن أن تكون ChatGPT وGemini مصادر موثوقة للمعلومات العامة حول التهاب الكبد الفيروسي، فإن التزامهما بالإرشادات الطبية لا يزال غير كافٍ. تسلط الدراسة الضوء على إمكانية هذه المنصات الذكية في تخفيف أعباء الرعاية الصحية من خلال إبلاغ الجمهور، لكنها تؤكد على الحاجة إلى تحسين موثوقيتها للمهنيين الصحيين. تشمل قيود الدراسة حجم العينة الصغيرة لأسئلة وسائل التواصل الاجتماعي والموضوعية المحتملة في التقييمات من قبل المتخصصين. بشكل عام، تشير الأبحاث إلى أنه مع المزيد من التحسين، يمكن أن تلعب هذه الروبوتات دورًا كبيرًا في تدابير الصحة الوقائية، على الرغم من أنها حاليًا لا تلبي المعايير في تقديم معلومات شاملة وموثوقة للاستخدام السريري.
DOI: https://doi.org/10.1038/s41598-024-83575-1
PMID: https://pubmed.ncbi.nlm.nih.gov/39799203
Publication Date: 2025-01-11
Author(s): Meryem Şahin Özdemir et al.
Primary Topic: Vaccine Coverage and Hesitancy
Overview
This study is the first to assess the adequacy and reliability of the ChatGPT and Gemini chatbots in providing information on viral hepatitis. A total of 176 questions were categorized into three groups: public Q&As from the CDC, strong recommendations from international guidelines, and frequently asked questions from social media. The responses from both chatbots were evaluated by infectious diseases specialists using a scoring scale, and inter-rater reliability was assessed using Cohen’s kappa coefficient. The results indicated that both ChatGPT and Gemini had similar mean scores (3.55 ± 0.83 vs. 3.57 ± 0.89, p = 0.260) and rates of completely correct answers (71.0% vs. 78.4%, p = 0.111). Subgroup analyses showed comparable performance across all question categories, although the accuracy for guideline-related questions was notably lower.
The study found a moderate positive correlation between the answers provided by ChatGPT and Gemini (r = 0.633, p < 0.001), with reproducibility rates of 91.3% for ChatGPT and 92% for Gemini (p = 0.710). Cohen's kappa values indicated substantial inter-rater agreement for both chatbots (κ = 0.720 for ChatGPT and κ = 0.704 for Gemini). While both chatbots effectively addressed CDC and social media questions, their performance on guideline questions was insufficient. The findings suggest that ChatGPT and Gemini could serve as valuable resources for public health information on viral hepatitis, but improvements are needed to enhance their reliability for healthcare professionals.
Methods
This study, conducted from March 1 to March 30, 2024, aimed to evaluate the quality of responses generated by ChatGPT and Google Gemini regarding viral hepatitis. The research involved three groups of questions: (1) “questions and answers (Q&As) for the public” from the Centers for Disease Control and Prevention (CDC), (2) evidence-based recommendations from international guidelines on viral hepatitis, and (3) frequently asked questions sourced from social media platforms and healthcare association websites. A total of 176 questions were included, with 34.7% from the CDC, 42.6% from guidelines, and 22.7% from social media. Rigorous criteria were applied to exclude irrelevant or unclear questions.
Responses from ChatGPT 3.5 and Google Gemini 1.5 Flash were evaluated by two infectious diseases and clinical microbiology specialists, with discrepancies resolved by a third expert. A standardized scoring system rated the answers on a scale from 1 (completely wrong) to 4 (completely correct and sufficient). To assess reproducibility, each question was posed to two different computers at separate times, categorizing responses as either consistent or inconsistent. The study did not require ethics committee approval as it did not involve patient data.
Results
The study evaluated the performance of two AI chatbots, ChatGPT and Gemini, in answering questions related to various types of hepatitis. A total of 176 questions were analyzed, with the distribution across hepatitis A (18.8%), B (29.5%), C (21.0%), D (17.0%), and E (13.7%). Inter-rater agreement was substantial for both chatbots, with Cohen’s kappa values of κ = 0.720 for ChatGPT and κ = 0.704 for Gemini. ChatGPT achieved a correct answer rate of 71.0%, while Gemini performed slightly better at 78.4%. Notably, both chatbots demonstrated higher accuracy in responding to CDC questions compared to guideline questions, with ChatGPT scoring 90.1% versus 49.4% and Gemini scoring 91.9% versus 61.4%.
Further analysis revealed that the accuracy of responses varied by hepatitis type, with the highest correct answer rates for hepatitis E (91.6%) and the lowest for hepatitis D (60.0%) for ChatGPT. For Gemini, the rates were similarly high for hepatitis E (91.6%) and varied for other types. The mean scores for the responses were comparable, with ChatGPT scoring 3.55 ± 0.83 and Gemini 3.57 ± 0.89, indicating no significant difference (p = 0.260). Additionally, the reproducibility of answers was high for both chatbots, at 91.3% for ChatGPT and 92% for Gemini, suggesting consistent performance across different question categories. Overall, while both chatbots performed well, Gemini exhibited a marginally higher accuracy rate.
Discussion
In this study, the reliability and knowledge of the ChatGPT and Gemini chatbots regarding viral hepatitis were evaluated using statistical analysis via SPSS version 25.0. Continuous variables were expressed as mean ± standard deviation, while categorical variables were presented as counts and percentages. The study found that both chatbots performed well in answering questions from the CDC (90.1% for ChatGPT and 91.9% for Gemini) and social media platforms (82.5% for ChatGPT and 90.0% for Gemini). However, their performance on guideline-based questions was notably lower, with correct answer rates of 49.4% for ChatGPT and 61.4% for Gemini. Despite high reproducibility rates (91.3% for ChatGPT and 92% for Gemini), both chatbots struggled with treatment recommendation questions, achieving only 42.8% accuracy.
The findings suggest that while ChatGPT and Gemini can serve as reliable sources for public information on viral hepatitis, their adherence to medical guidelines remains inadequate. The study highlights the potential of these AI platforms to alleviate healthcare burdens by informing the public, yet emphasizes the need for improvement in their reliability for healthcare professionals. Limitations of the study include a small sample size for social media questions and potential subjectivity in evaluations by specialists. Overall, the research indicates that with further refinement, these chatbots could play a significant role in preventive health measures, although they currently fall short in providing comprehensive and reliable information for clinical use.
