DOI: https://doi.org/10.1007/s00432-026-06432-w
PMID: https://pubmed.ncbi.nlm.nih.gov/41758476
تاريخ النشر: 2026-02-27
المؤلف: Reza Dehdab وآخرون
الموضوع الرئيسي: الذكاء الاصطناعي في الرعاية الصحية والتعليم
نظرة عامة
تستكشف هذه الدراسة إمكانيات ChatGPT-4o، وهو نموذج لغوي كبير، لتعزيز سير العمل في مجال مجالس الأورام متعددة التخصصات (MDTs) في إدارة ساركومات الأنسجة الرخوة (STS). نظرًا للقيود الزمنية والتكلفة والموارد في MDTs، كانت الأبحاث تهدف إلى تقييم الأداء السريري لـ ChatGPT-4o من خلال مقارنة توصيات العلاج التي يقدمها مع تلك التي تقدمها MDTs الخبراء. تم تحليل 152 حالة مريض، حيث قام ChatGPT-4o بتوليد اقتراحات علاجية مستندة إلى الإرشادات من رسائل تسجيل مجالس الأورام المجهولة. تم تقييم النتائج من قبل مراجعين خبراء معصوبي العيون عبر خمسة مجالات: طرق التشخيص، طرق العلاج، تسلسل/توقيت العلاج، نظام العلاج الكيميائي، والسياق السريري.
أشارت النتائج إلى أن درجات ChatGPT-4o كانت أقل بكثير من القيمة القصوى الممكنة 1.0 في جميع المعايير التي تم تقييمها (جميعها p < 0.0001). ومن الجدير بالذكر أن السياق السريري حصل على أعلى الدرجات مقارنة بالمجالات الأخرى (جميعها p < 0.05)، بينما لم يتم العثور على اختلافات كبيرة عبر أنواع الساركومات المختلفة (H = 19.74، p = 0.138). تشير النتائج إلى أنه بينما يظهر ChatGPT-4o وعدًا في توليد التوصيات لإدارة STS، وخاصة في السياق الشخصي، فإن الفجوات في تسلسل العلاج واختيار العلاج الكيميائي تؤكد على ضرورة الإشراف من قبل الخبراء. تدعم هذه الأبحاث دمج النماذج اللغوية الكبيرة في سير العمل في علم الأورام، مما يبرز الحاجة إلى مزيد من التحسين لضمان تطبيق سريري آمن وفعال.
مقدمة
تُعتبر ساركومات الأنسجة الرخوة (STS) مجموعة نادرة ومتنوعة من الأورام الخبيثة، تمثل حوالي 1% من جميع سرطانات البالغين. تتطلب إدارتها نهجًا متعدد التخصصات، يشمل التعاون بين تخصصات مختلفة مثل الجراحة، والطب، وعلم الأورام الإشعاعي، بالإضافة إلى علم الأمراض والأشعة. تلعب مجالس الأورام متعددة التخصصات (MDT) دورًا حاسمًا في هذه العملية من خلال تسهيل تقييمات شاملة للحالات واستراتيجيات علاج شخصية. ومع ذلك، يمكن أن تعيق التحديات اللوجستية – بما في ذلك القيود المالية، ومتطلبات الوقت، والفجوات الجغرافية – التنفيذ الفعال لـ MDTs. وقد أدى ذلك إلى زيادة الاهتمام في الاستفادة من الذكاء الاصطناعي (AI) لتعزيز اتخاذ القرار داخل هذه الفرق.
أنتجت التقدمات الأخيرة في الذكاء الاصطناعي، وخاصة في معالجة اللغة الطبيعية (NLP)، نماذج لغوية كبيرة (LLMs) مثل سلسلة المحولات المدربة مسبقًا (GPT)، التي يمكنها توليد نصوص شبيهة بالبشر وأداء مهام متنوعة مع تدريب محدود. بينما استكشفت الدراسات الأولية استخدام الذكاء الاصطناعي في مجالس الأورام لأنواع مختلفة من السرطان، لا يزال تطبيق LLMs في سياق الساركومات غير مُفحص إلى حد كبير. تهدف هذه الدراسة إلى تقييم أداء ChatGPT-4o في حالات الساركومات الواقعية مقابل معايير سريرية محددة مسبقًا، وتقييم توافق اقتراحاته العلاجية مع توصيات الخبراء من فرق الأورام متعددة التخصصات ذات الخبرة. يمكن أن توضح النتائج الدور المحتمل لـ LLMs كأدوات داعمة في علم الأورام، وخاصة في تخطيط العلاج واتخاذ القرار، وتوجيه سير العمل السريري المدعوم بالذكاء الاصطناعي في المستقبل.
الطرق
توضح قسم “المواد والطرق” تصميم التجربة والإجراءات المستخدمة في الدراسة. يتفصيل المواد المحددة المستخدمة، بما في ذلك أي مواد كيميائية، ومعدات، وعينات بيولوجية، لضمان إمكانية تكرار البحث. تشمل المنهجية البروتوكولات المتبعة لجمع البيانات، بما في ذلك أي تحليلات إحصائية تم تطبيقها لتفسير النتائج.
بالإضافة إلى ذلك، قد يصف القسم الظروف التجريبية، مثل درجة الحرارة، والمدة، وأي ضوابط تم تنفيذها للتحقق من النتائج. بشكل عام، يخدم هذا القسم لتقديم نظرة شاملة على التقنيات والعمليات التي تدعم البحث، مما يسمح بالتقييم النقدي والتكرار من قبل باحثين آخرين في هذا المجال.
النتائج
يقدم قسم “النتائج” من ورقة البحث النتائج الرئيسية المستمدة من التجارب أو التحليلات التي تم إجراؤها. يبرز نقاط البيانات المهمة، والاتجاهات، والنتائج الإحصائية التي تدعم الفرضيات أو أسئلة البحث المطروحة سابقًا في الدراسة. غالبًا ما تكون النتائج مصحوبة بوسائل بصرية مثل الرسوم البيانية أو الجداول لتسهيل الفهم وتوضيح العلاقات بين المتغيرات.
في هذا القسم، قد يناقش المؤلفون أيضًا تداعيات نتائجهم، مقارنتها بالأدبيات الموجودة لوضع مساهماتهم في السياق. يتم تناول أي نتائج غير متوقعة أو شذوذ، مما يوفر رؤى حول القيود المحتملة أو المجالات التي تحتاج إلى مزيد من التحقيق. بشكل عام، تعتبر النتائج أساسًا للنقاشات والاستنتاجات اللاحقة التي تم التوصل إليها في الورقة.
المناقشة
هدفت الدراسة التي أجريت في مستشفى جامعة توبنغن إلى تقييم أداء ChatGPT-4o في توليد توصيات العلاج للمرضى الذين يعانون من ساركومات الأنسجة الرخوة (STS) بما يتماشى مع قرارات مجالس الأورام متعددة التخصصات (MTB). باستخدام تصميم استعادي، شملت الأبحاث مجموعة من 138 مريضًا، مع تطوير نظام تقييم لتقييم جودة الاقتراحات المعتمدة على LLM مقابل معايير اتخاذ القرار السريرية المعتمدة. كشفت النتائج عن مستوى عالٍ من التوافق بين توصيات النموذج وتوصيات MTB، مع تميز خاص في السياق السريري، مما يبرز قدرة النموذج على دمج العوامل الخاصة بالمرضى. ومع ذلك، تم ملاحظة قيود في مجالات طرق العلاج، تسلسل العلاج، واختيار نظام العلاج الكيميائي، مما يشير إلى أنه بينما يمكن للنموذج المساعدة في تجميع المعلومات السريرية، فإنه يواجه صعوبة في توليد مسارات علاج دقيقة.
على الرغم من إمكانياته كأداة داعمة في سير العمل في علم الأورام، تؤكد الدراسة على ضرورة الإشراف البشري في اتخاذ القرارات السريرية. يبرز المؤلفون أنه بينما أظهر ChatGPT-4o أداءً كبيرًا، خاصة في التفكير السياقي، فإن قيوده في توليد خطط علاج تفصيلية تتطلب مزيدًا من التطوير والتحقق. كما تعترف الدراسة بعدة قيود، بما في ذلك تصميمها في مركز واحد والتحديات المرتبطة باستخدام نموذج مغلق المصدر، مما قد يؤثر على إمكانية التكرار وخصوصية البيانات. بشكل عام، تقترح الأبحاث أن LLMs مثل ChatGPT-4o يمكن أن تعزز كفاءة مناقشات مجالس الأورام ولكن لا ينبغي أن تحل محل حكم الخبراء في الممارسة السريرية.
DOI: https://doi.org/10.1007/s00432-026-06432-w
PMID: https://pubmed.ncbi.nlm.nih.gov/41758476
Publication Date: 2026-02-27
Author(s): Reza Dehdab et al.
Primary Topic: Artificial Intelligence in Healthcare and Education
Overview
This study investigates the potential of ChatGPT-4o, a large language model, to enhance the workflows of multidisciplinary tumor boards (MDTs) in the management of soft tissue sarcomas (STS). Given the constraints of time, cost, and resources in MDTs, the research aimed to evaluate the clinical performance of ChatGPT-4o by comparing its treatment recommendations against those made by expert MDTs. A total of 152 patient cases were analyzed, with ChatGPT-4o generating guideline-based treatment suggestions from anonymized tumor board registration letters. The outputs were assessed by blinded expert reviewers across five domains: diagnostic modalities, therapeutic modalities, treatment sequencing/timing, chemotherapy regimen, and clinical contextualization.
The results indicated that ChatGPT-4o’s scores were significantly lower than the maximum achievable value of 1.0 in all evaluated criteria (all p < 0.0001). Notably, clinical contextualization received the highest scores compared to other domains (all p < 0.05), while no significant differences were found across various sarcoma subtypes (H = 19.74, p = 0.138). The findings suggest that while ChatGPT-4o shows promise in generating recommendations for STS management, particularly in personalized contextualization, discrepancies in treatment sequencing and chemotherapy selection underscore the necessity for expert oversight. This research supports the integration of large language models into oncology workflows, highlighting the need for further refinement to ensure safe and effective clinical application.
Introduction
Soft tissue sarcomas (STS) are a rare and diverse group of malignancies, accounting for about 1% of all adult cancers. Their management necessitates a multidisciplinary approach, involving collaboration among various specialties such as surgical, medical, and radiation oncology, as well as pathology and radiology. Multidisciplinary tumor boards (MDT) play a crucial role in this process by facilitating comprehensive case assessments and personalized treatment strategies. However, logistical challenges—including financial constraints, time demands, and geographic disparities—can hinder the effective implementation of MDTs. This has led to an increasing interest in leveraging artificial intelligence (AI) to enhance decision-making within these teams.
Recent advancements in AI, particularly in natural language processing (NLP), have produced large language models (LLMs) like the Generative Pretrained Transformer (GPT) series, which can generate human-like text and perform various tasks with minimal training. While initial studies have explored the use of AI in tumor boards for different cancer types, the application of LLMs in the context of sarcomas remains largely unexamined. This study aims to evaluate the performance of ChatGPT-4o in real-world sarcoma cases against predefined clinical criteria, assessing the alignment of its treatment suggestions with expert recommendations from experienced multidisciplinary oncology teams. The findings could elucidate the potential role of LLMs as supportive tools in oncology, particularly in treatment planning and decision-making, and inform future AI-assisted clinical workflows.
Methods
The “Materials and Methods” section outlines the experimental design and procedures employed in the study. It details the specific materials used, including any reagents, equipment, and biological samples, ensuring reproducibility of the research. The methodology encompasses the protocols followed for data collection, including any statistical analyses applied to interpret the results.
Additionally, the section may describe the experimental conditions, such as temperature, duration, and any controls implemented to validate the findings. Overall, this section serves to provide a comprehensive overview of the techniques and processes that underpin the research, allowing for critical evaluation and replication by other scholars in the field.
Results
The “Results” section of the research paper presents key findings derived from the conducted experiments or analyses. It highlights significant data points, trends, and statistical outcomes that support the hypotheses or research questions posed earlier in the study. The results are often accompanied by visual aids such as graphs or tables to facilitate comprehension and illustrate the relationships between variables.
In this section, the authors may also discuss the implications of their findings, comparing them with existing literature to contextualize their contributions to the field. Any unexpected results or anomalies are addressed, providing insights into potential limitations or areas for further investigation. Overall, the results serve as a foundation for the subsequent discussion and conclusions drawn in the paper.
Discussion
The study conducted at the University Hospital Tübingen aimed to evaluate the performance of ChatGPT-4o in generating treatment recommendations for patients with soft tissue sarcomas (STS) in alignment with multidisciplinary tumor board (MTB) decisions. Utilizing a retrospective design, the research involved a cohort of 138 patients, with a scoring system developed to assess the quality of LLM-based suggestions against established clinical decision criteria. The findings revealed a high level of concordance between the model’s recommendations and those of the MTB, particularly excelling in clinical contextualization, which emphasizes the model’s ability to integrate patient-specific factors. However, limitations were noted in the areas of therapeutic modalities, treatment sequencing, and chemotherapy regimen selection, indicating that while the model can assist in synthesizing clinical information, it struggles with generating precise treatment pathways.
Despite its potential as a supportive tool in oncology workflows, the study underscores the necessity of human oversight in clinical decision-making. The authors highlight that while ChatGPT-4o showed substantial performance, particularly in contextual reasoning, its limitations in generating detailed therapeutic plans necessitate further development and validation. The study also acknowledges several limitations, including its single-center design and the challenges associated with using a closed-source model, which may impact reproducibility and data privacy. Overall, the research suggests that LLMs like ChatGPT-4o could enhance the efficiency of tumor board discussions but should not replace expert judgment in clinical practice.
