تقييم نماذج اللغة الكبيرة متعددة الوسائط لتشخيص الصدفية، وتصنيف شدة المرض، وتوصيات العلاج من الصور السريرية: ChatGPT يظهر أداءً متفوقًا مقارنةً بنماذج اللغة الكبيرة الأخرى
Evaluation of multimodal large language models for psoriasis diagnosis, severity grading, and treatment recommendations from clinical photographs: ChatGPT shows superior performance compared to other large language models

شارك:
المجلة: Frontiers in Medicine، المجلد: 13
DOI: https://doi.org/10.3389/fmed.2026.1791488
PMID: https://pubmed.ncbi.nlm.nih.gov/42210965
تاريخ النشر: 2026-05-13
المؤلف: Mehdi Boostani وآخرون
الموضوع الرئيسي: الصدفية: العلاج والآلية المرضية

نظرة عامة

تقوم الدراسة بتقييم فعالية نماذج اللغة الكبيرة متعددة الوسائط (LLMs) المستندة إلى الويب في تشخيص الصدفية، وتقييم شدتها من خلال تقييم الطبيب العالمي (PGA)، وتقديم توصيات العلاج استنادًا إلى الصور السريرية. تم إجراء تحليل استعادي على 303 صور موحدة من 160 مريضًا، والتي شملت كل من آفات الصدفية وحالات التشابه. وجدت الدراسة أن ChatGPT-5 وChatGPT-4o أظهرا أعلى دقة تشخيصية بنسبة 93.1% و90.1%، على التوالي، مع كون دقة PGA أيضًا مرتفعة بشكل ملحوظ لهذه النماذج. كانت توصيات العلاج الأكثر ملاءمة من ChatGPT-5 (83.7%) وChatGPT-4o (82.1%).

تؤكد الخاتمة على أنه بينما تظهر هذه LLMs وعدًا في دعم التعرف على الصدفية وتقييمها، إلا أنها ليست مناسبة بعد للاستخدام المستقل من قبل المرضى. تشير النتائج إلى أن ChatGPT-5 وChatGPT-4o يمكن أن تكون أدوات قيمة للأطباء، خاصة في إعدادات الرعاية الأولية حيث يكون الوصول إلى المتخصصين في الأمراض الجلدية محدودًا. ومع ذلك، يحذر المؤلفون من استخلاص استنتاجات أوسع بشأن فائدتها السريرية حتى يتم إجراء مزيد من التحقق في مجموعات بيانات واقعية متعددة المراكز ومتنوعة.

مقدمة

تتناول مقدمة ورقة البحث الصدفية، وهي حالة جلدية التهابية مزمنة ناتجة عن المناعة تؤثر على 0.91 إلى 8.5% من البالغين على مستوى العالم. يبرز المؤلفون التحديات في تشخيص الصدفية في الرعاية الأولية بسبب تقديمها المتغير وتشابهها مع حالات جلدية أخرى، مثل الإكزيما والسعفة. تتفاقم هذه الصعوبة التشخيصية بسبب عدم التعرف عليها في الأفراد ذوي درجات لون البشرة الداكنة وفي المواقع غير النمطية، مما يؤدي إلى تأخيرات في العلاج قد تؤدي إلى أعراض مستمرة ومشاكل نفسية مرتبطة، بما في ذلك القلق والاكتئاب. يتم التأكيد على أهمية التشخيص المبكر والدقيق، حيث إنه أمر حاسم لبدء العلاجات المستندة إلى الأدلة وتقليل النتائج السلبية.

تناقش المقدمة أيضًا دور تقييم الطبيب العالمي (PGA) في تقييم شدة المرض واستجابة العلاج، حيث يعمل كأداة حاسمة لاتخاذ القرارات السريرية. علاوة على ذلك، تقدم الورقة نماذج اللغة الكبيرة (LLMs)، وخاصة LLMs متعددة الوسائط التي يمكنها معالجة كل من الصور والنصوص، كمسار واعد لتعزيز التشخيصات الجلدية. تهدف الدراسة إلى تقييم أداء هذه LLMs المستندة إلى الويب في تشخيص الصدفية، وتسجيل PGA، وتقديم توصيات العلاج استنادًا إلى الصور السريرية، مما يبرز أهمية هذا البحث في تحسين الرعاية الجلدية.

الطرق

تحدد قسم “الطرق” في ورقة البحث التصميم التجريبي والتقنيات التحليلية المستخدمة للتحقيق في أسئلة البحث. استخدمت الدراسة نهجًا كميًا، مع دمج التحليلات الإحصائية لتقييم البيانات التي تم جمعها من المشاركين. شملت المنهجيات المحددة تجارب محكومة، واستطلاعات، ودراسات رصدية، مما يضمن إطارًا قويًا لجمع البيانات.

تم إجراء تحليل البيانات باستخدام برامج إحصائية قياسية، مع التركيز على نماذج الانحدار لتحديد العلاقات بين المتغيرات. تأكد الباحثون من صحة وموثوقية نتائجهم من خلال الاختبارات الدقيقة والالتزام بالإرشادات الأخلاقية. بشكل عام، تم تصميم الطرق المستخدمة لتوفير فهم شامل للظواهر قيد التحقيق، مما يسهل تفسير النتائج في سياق الأدبيات الموجودة.

النتائج

يقدم قسم “النتائج” في ورقة البحث النتائج الرئيسية المستمدة من التجارب أو التحليلات التي تم إجراؤها. عادةً ما يتضمن بيانات كمية، وتحليلات إحصائية، وتمثيلات بصرية مثل الرسوم البيانية أو الجداول لتوضيح النتائج. غالبًا ما تتم مقارنة النتائج مع الفرضيات أو الدراسات السابقة لتسليط الضوء على الفروق أو التأكيدات المهمة.

في هذا القسم، قد يبلغ المؤلفون عن مقاييس محددة، مثل المتوسطات، والانحرافات المعيارية، وقيم p، لدعم ادعاءاتهم. بالإضافة إلى ذلك، تتم مناقشة أي اتجاهات أو أنماط ملحوظة، مما يوفر رؤى حول تداعيات النتائج. بشكل عام، تعتبر النتائج أساسًا للنقاشات اللاحقة والاستنتاجات المستخلصة في الورقة.

النقاش

قيمت الدراسة القدرات التشخيصية لأربعة نماذج لغة كبيرة متعددة الوسائط (LLMs)—ChatGPT-4o، ChatGPT-5، Gemini 2.5 Flash، وClaude Sonnet 4.5—باستخدام صور سريرية للصدفية وحالات تشبه الصدفية. أجريت في جامعة سملويس، شمل التحليل الاستعادي 303 صورة، مع التركيز على الدقة التشخيصية، وتسجيل تقييم الطبيب العالمي (PGA)، وملاءمة توصيات العلاج. برز ChatGPT-5 كأفضل أداء، محققًا دقة تشخيصية بنسبة 93.1%، بينما جاء ChatGPT-4o قريبًا عند 90.1%. أظهرت كلا النموذجين دقة عالية في تسجيل PGA وملاءمة توصيات العلاج، على الرغم من أنهما لم يصلوا إلى الدقة التشخيصية المثالية للأطباء الجلدين المشاركين في الدراسة.

تسلط النتائج الضوء على إمكانيات LLMs كأدوات داعمة في الأمراض الجلدية، خاصة في تصنيف حالات الصدفية. ومع ذلك، تظل أداء النماذج، على الرغم من وعدها، أقل من أداء الأطباء الجلدين ذوي الخبرة، مما يشير إلى الحاجة إلى مزيد من التحقق في إعدادات واقعية متنوعة. تؤكد الدراسة على أهمية الإشراف البشري والاعتبارات الأخلاقية في نشر تقنيات الذكاء الاصطناعي في الممارسة السريرية، خاصة فيما يتعلق بمخاطر المخرجات غير الصحيحة وضرورة وجود أطر حوكمة قوية. يجب أن تهدف الأبحاث المستقبلية إلى تضمين مجموعات بيانات أكبر ومتعددة المراكز ومعالجة تباين ظروف التصوير في العالم الحقيقي لتعزيز قابلية تعميم هذه النتائج.

القيود

تُعزى قيود هذه الدراسة بشكل أساسي إلى تصميمها المتمركز في مركز واحد واستخدام مجموعة صور موحدة تتميز بشكل أساسي بأنواع جلد فيتزباتريك II-III، والتي قد لا تمثل بشكل كاف تنوع تقديمات الصدفية عبر درجات لون البشرة المختلفة. تتكون مجموعة البيانات بشكل رئيسي من صدفية من نوع اللويحات، مع تمثيل غير كاف لأنواع سريرية أخرى، مما قد يؤثر على قابلية تعميم النتائج. بالإضافة إلى ذلك، لم تتضمن الدراسة مقارنة رسمية مع مصنفات الشبكات العصبية التلافيفية (CNN) المعتمدة، مما يحد من تقييم الأداء النسبي لنماذج اللغة الكبيرة (LLMs) مقابل نماذج الذكاء الاصطناعي المرجعية.

علاوة على ذلك، يتطلب التطور السريع لنماذج LLMs المستندة إلى الويب إعادة تقييم مستمرة وإشراف بشري. كما أن غياب خوارزميات تشخيص موحدة للصدفية يعقد عملية القياس، حيث تعتمد التقييمات عادةً على تقييمات الأطباء الجلدين أو نماذج CNN الخاصة بالمهام التي لم يتم تضمينها في هذا البحث. قد تكون الظروف المحكومة التي تم التقاط الصور فيها، بينما تعزز التناسق الداخلي، قد أدت إلى إدخال تحيز في الاكتساب، مما أدى إلى مجموعة صور أقل تمثيلًا مقارنة بالسيناريوهات السريرية في العالم الحقيقي. وبالتالي، قد لا يكون أداء النموذج قابلاً للتطبيق على الفئات السكانية المرضية الممثلة تمثيلًا ناقصًا، خاصةً أولئك ذوي درجات لون البشرة الداكنة، حيث قد تظهر الصدفية بشكل مختلف.

Journal: Frontiers in Medicine, Volume: 13
DOI: https://doi.org/10.3389/fmed.2026.1791488
PMID: https://pubmed.ncbi.nlm.nih.gov/42210965
Publication Date: 2026-05-13
Author(s): Mehdi Boostani et al.
Primary Topic: Psoriasis: Treatment and Pathogenesis

Overview

The research evaluates the effectiveness of web-based multimodal large language models (LLMs) in diagnosing psoriasis, assessing its severity through Physician Global Assessment (PGA) scoring, and providing treatment recommendations based on clinical photographs. A retrospective analysis was conducted on 303 standardized images from 160 patients, which included both psoriasis lesions and mimicking conditions. The study found that ChatGPT-5 and ChatGPT-4o exhibited the highest diagnostic accuracy at 93.1% and 90.1%, respectively, with PGA accuracy also being notably high for these models. Treatment recommendations were most appropriate from ChatGPT-5 (83.7%) and ChatGPT-4o (82.1%).

The conclusion emphasizes that while these LLMs show promise in supporting psoriasis recognition and assessment, they are not yet suitable for independent patient use. The findings suggest that ChatGPT-5 and ChatGPT-4o could serve as valuable tools for clinicians, particularly in primary care settings where access to dermatological specialists is limited. However, the authors caution against drawing broader conclusions regarding their clinical utility until further validation is conducted in diverse, multi-center real-world datasets.

Introduction

The introduction of the research paper addresses psoriasis, a chronic immune-mediated inflammatory skin condition affecting 0.91 to 8.5% of adults globally. The authors highlight the challenges in diagnosing psoriasis in primary care due to its variable presentation and similarities with other skin conditions, such as eczema and tinea. This diagnostic difficulty is exacerbated by under-recognition in individuals with darker skin tones and in atypical locations, leading to delays in treatment that can result in persistent symptoms and associated psychological issues, including anxiety and depression. The importance of early and accurate diagnosis is emphasized, as it is crucial for initiating evidence-based therapies and minimizing adverse outcomes.

The introduction also discusses the role of the Physician Global Assessment (PGA) in evaluating disease severity and treatment response, serving as a critical tool for clinical decision-making. Furthermore, the paper introduces large language models (LLMs), particularly multimodal LLMs that can process both images and text, as a promising avenue for enhancing dermatological diagnostics. The study aims to evaluate the performance of these web-based LLMs in diagnosing psoriasis, scoring PGA, and providing treatment recommendations based on clinical photographs, highlighting the relevance of this research in improving dermatological care.

Methods

The “Methods” section of the research paper outlines the experimental design and analytical techniques employed to investigate the research questions. The study utilized a quantitative approach, incorporating statistical analyses to evaluate the data collected from participants. Specific methodologies included controlled experiments, surveys, and observational studies, ensuring a robust framework for data collection.

Data analysis was performed using standard statistical software, with emphasis on regression models to identify relationships between variables. The researchers ensured the validity and reliability of their findings through rigorous testing and adherence to ethical guidelines. Overall, the methods employed were designed to provide a comprehensive understanding of the phenomena under investigation, facilitating the interpretation of results in the context of existing literature.

Results

The “Results” section of the research paper presents the key findings derived from the conducted experiments or analyses. It typically includes quantitative data, statistical analyses, and visual representations such as graphs or tables to illustrate the outcomes. The results are often compared against the hypotheses or previous studies to highlight significant differences or confirmations.

In this section, the authors may report specific metrics, such as means, standard deviations, and p-values, to substantiate their claims. Additionally, any observed trends or patterns are discussed, providing insights into the implications of the findings. Overall, the results serve as a foundation for the subsequent discussion and conclusions drawn in the paper.

Discussion

The study evaluated the diagnostic capabilities of four multimodal large language models (LLMs)—ChatGPT-4o, ChatGPT-5, Gemini 2.5 Flash, and Claude Sonnet 4.5—using clinical photographs of psoriasis and psoriasis-mimicking conditions. Conducted at Semmelweis University, the retrospective analysis involved 303 images, with a focus on diagnostic accuracy, Physician Global Assessment (PGA) scoring, and treatment recommendation appropriateness. ChatGPT-5 emerged as the top performer, achieving a diagnostic accuracy of 93.1%, while ChatGPT-4o followed closely at 90.1%. Both models demonstrated high PGA scoring accuracy and treatment recommendation appropriateness, although they did not reach the perfect diagnostic accuracy of the dermatologists involved in the study.

The findings highlight the potential of LLMs as supportive tools in dermatology, particularly in triaging psoriasis cases. However, the models’ performance, while promising, remains below that of experienced dermatologists, indicating the need for further validation in diverse, real-world settings. The study emphasizes the importance of human oversight and ethical considerations in deploying AI technologies in clinical practice, particularly regarding the risks of incorrect outputs and the necessity for robust governance frameworks. Future research should aim to include larger, multi-center datasets and address the variability of real-world imaging conditions to enhance the generalizability of these findings.

Limitations

The limitations of this study are primarily attributed to its single-center design and the use of a standardized image set that predominantly features Fitzpatrick skin types II-III, which may not adequately represent the diversity of psoriasis presentations across different skin tones. The dataset mainly consists of plaque-type psoriasis, with other clinical subtypes underrepresented, potentially affecting the generalizability of the findings. Additionally, the study did not include a formal comparison with established convolutional neural network (CNN)-based classifiers, limiting the assessment of the relative performance of large language models (LLMs) against benchmark AI models.

Moreover, the rapid evolution of web-based LLMs necessitates ongoing reevaluation and human oversight. The absence of standardized diagnostic algorithms for psoriasis further complicates benchmarking, as assessments typically rely on dermatologist evaluations or task-specific CNN models not included in this research. The controlled conditions under which images were captured, while enhancing internal consistency, may have introduced acquisition bias, resulting in a less representative image set compared to real-world clinical scenarios. Consequently, the model’s performance may not be applicable to underrepresented patient populations, particularly those with darker skin tones, where psoriasis may manifest differently.

شارك: