DOI: https://doi.org/10.1038/s43856-025-00808-8
PMID: https://pubmed.ncbi.nlm.nih.gov/40164789
تاريخ النشر: 2025-03-31
المؤلف: Bastian Grothey وآخرون
الموضوع الرئيسي: الذكاء الاصطناعي في الرعاية الصحية والتعليم
نظرة عامة
تبحث الدراسة في استخراج البيانات المنظمة من تقارير الأمراض غير المنظمة باستخدام نماذج لغوية مختلفة، مع تسليط الضوء على التحديات والإمكانات لكل من الحلول المملوكة والمفتوحة المصدر. تم إنشاء مجموعة بيانات تضم 579 تقريرًا مرضيًا مشروحًا بالألمانية والإنجليزية لتقييم أداء ستة نماذج لغوية، بما في ذلك نموذج GPT-4 المملوك وعدة نماذج مفتوحة المصدر (عائلات Llama2 وLlama3، Qwen2.5). تؤكد الدراسة على أهمية هندسة المطالبات وتقنيات تقليل حجم النموذج للنشر الفعال على الأجهزة الاستهلاكية، كاشفة أن النماذج المفتوحة المصدر يمكن أن تحقق مستويات دقة مقارنة بـ GPT-4.
تشير النتائج إلى أن نماذج اللغة المفتوحة المصدر لا توفر فقط بديلاً فعالاً من حيث التكلفة ويحافظ على الخصوصية لهيكلة بيانات الأمراض، ولكنها أيضًا تظهر تباينًا كبيرًا في الأداء بناءً على التكوينات المستخدمة. تؤكد الدراسة على الإمكانية لهذه النماذج لتسهيل تطوير أدوات الذكاء الاصطناعي في الرعاية الصحية، خاصة في علم الأورام، حيث يتم أرشفة كميات هائلة من البيانات القيمة. تعتبر مجموعة البيانات الثنائية المتاحة للجمهور معيارًا للبحوث المستقبلية، تهدف إلى تعزيز الاستخراج المنهجي ووضع العلامات على البيانات اللازمة لتدريب خوارزميات الذكاء الاصطناعي في المجال الطبي.
طرق
في هذه الدراسة، تم تحليل 579 تقريرًا مرضيًا من 340 مريضًا مصابًا بسرطان البروستاتا الذين خضعوا لاستئصال البروستاتا الجذري بين عامي 2020 و2022 لتقييم أداء نماذج لغوية مختلفة في استخراج البيانات المنظمة. تم استرجاع التقارير بتنسيق نص عادي مجهول الهوية، وتم توضيحها بناءً على تصنيفات معتمدة، بما في ذلك تصنيف TNM للأورام الخبيثة ونظام تصنيف جليسون. تم اختيار أحد عشر معلمًا سريريًا ذي صلة للاستخراج، بما في ذلك مجموعة درجة WHO (ISUP)، مرحلة T، ومشاركة العقد اللمفاوية، والتي تعتبر حاسمة لقرارات التصنيف والعلاج. تم ترجمة التقارير من الألمانية إلى الإنجليزية باستخدام واجهة برمجة التطبيقات DeepL، والتزمت الدراسة بالإرشادات الأخلاقية، مع الحصول على موافقة المرضى لاستخدام البيانات المجهولة.
شمل التقييم مقارنة قدرات استخراج البيانات المنظمة لنماذج GPT-4 وLlama2/Llama3 تحت ظروف مختلفة. تم تكليف النماذج باستخراج المعلمات المحددة وتقديم المخرجات بتنسيق JSON منظم. تم حساب مقاييس الأداء، بما في ذلك الدقة، والموثوقية، والاسترجاع، ودرجة F1، لتقييم فعالية النماذج. استكشفت الدراسة أيضًا جدوى استخدام النماذج المقلصة للوصول الأوسع، مما يسمح باستخراج البيانات المنظمة على الأجهزة الاستهلاكية. بشكل عام، تسلط هذه الدراسة الضوء على إمكانيات نماذج اللغة في أتمتة استخراج المعلومات الحيوية من تقارير الأمراض، وهو أمر ضروري لتقدم تطبيقات الذكاء الاصطناعي في السياقات الطبية.
نتائج
أظهر تقييم نماذج اللغة الكبيرة المفتوحة المصدر والمملوكة (LLMs) أن Llama3، وخاصة النسخة 70B، تحقق دقة مقارنة بـ GPT-4 في استخراج البيانات المنظمة من تقارير الأمراض. حقق كلا النموذجين دقة إجمالية تزيد عن 97% عبر أحد عشر معلمة باللغتين الإنجليزية والألمانية. كما أن نموذج Llama3 الأصغر 8B قدم أداءً جيدًا، محققًا دقة 91% و83% في اللغتين المعنيتين. في المقابل، أظهر الجيل السابق، Llama2، أداءً أقل بكثير بغض النظر عن حجم النموذج. كشفت تحليلات الدقة على مستوى المعلمة أن تصنيف الأورام ونتائج الجراحة كانت الأكثر دقة في الاستخراج، بينما شكلت المعلمات الكمية تحديات أكبر، خاصة بالنسبة لـ Llama2.
فيما يتعلق بتوليد مخرجات JSON، أنتج GPT-4 ملفات JSON صالحة لجميع التقارير، بينما أظهر Llama3 70B نتائج شبه مثالية. واجه Llama2 صعوبات في تنسيق JSON، مما أدى غالبًا إلى مخرجات غير صالحة. من الجدير بالذكر أن جزءًا كبيرًا من ملفات JSON غير الصالحة اختلف بين الجولات الأولية وإعادة الاختبار، مما يشير إلى مشاكل جوهرية في توليد JSON بدلاً من مشاكل إدخال محددة. بالإضافة إلى ذلك، اختلفت أوقات المعالجة بشكل كبير، حيث بلغ متوسط GPT-4 12 ثانية لكل تقرير، بينما كان Llama3 8B أسرع بشكل ملحوظ عند 5-6 ثوانٍ. ومع ذلك، استغرق Llama2 70B وقتًا أطول بكثير، بمتوسط 51 ثانية. بشكل عام، تسلط الدراسة الضوء على الأداء المتفوق والكفاءة لـ Llama3 مقارنة بسابقه، Llama2، وتضع كل من Llama3 وGPT-4 كأدوات فعالة لاستخراج البيانات في مجال الأمراض.
مناقشة
في هذه الدراسة، أجرى المؤلفون تقييمًا شاملاً لمختلف نماذج اللغة الكبيرة (LLMs) لاستخراج البيانات المنظمة من تقارير الأمراض الجراحية، مع التركيز بشكل خاص على حالات استئصال البروستاتا الجذري. أنشأوا مجموعة بيانات عالية الجودة تحتوي على معلومات صحيحة واختبروا ستة نماذج LLM مختلفة، بما في ذلك النماذج المملوكة (GPT-4) والمفتوحة المصدر (Llama2، Llama3، Qwen2.5). أظهرت النتائج أنه بينما حقق GPT-4 دقة إجمالية مثيرة للإعجاب بلغت 98% باللغة الإنجليزية و97% بالألمانية، فإن النماذج المفتوحة المصدر مثل Llama3 قدمت أيضًا أداءً جيدًا، محققة مستويات دقة مماثلة دون التكاليف المرتبطة ومخاوف الخصوصية للنماذج المملوكة.
سلط التحليل الضوء على أهمية هندسة المطالبات وتقنيات التكميم في تحسين أداء النموذج. على الرغم من أن النسخ المقلصة من LLMs (4 بت) أظهرت انخفاضًا في الدقة مقارنة بنظيراتها كاملة الوزن، فإن استخدام استراتيجيات المطالبات الفعالة، وخاصة نهج سلسلة التحقق، حسن النتائج. وجدت الدراسة أن الهلاوس – حالات توليد معلومات غير موجودة – كانت أكثر شيوعًا في نماذج Llama2، بينما أظهر GPT-4 الحد الأدنى من الهلاوس. بالإضافة إلى ذلك، أثر تعقيد النص على دقة الاستخراج، مع وجود ارتباط ملحوظ بين تنوع المفردات والأداء. بشكل عام، تشير النتائج إلى أن LLMs المفتوحة المصدر يمكن أن تنظم بيانات تقارير الأمراض بشكل فعال، مما يوفر مسارًا واعدًا لتطبيقات الذكاء الاصطناعي في المجال الطبي، على الرغم من وجود مجال للتحسين في أداء النماذج المقلصة.
DOI: https://doi.org/10.1038/s43856-025-00808-8
PMID: https://pubmed.ncbi.nlm.nih.gov/40164789
Publication Date: 2025-03-31
Author(s): Bastian Grothey et al.
Primary Topic: Artificial Intelligence in Healthcare and Education
Overview
The research investigates the extraction of structured data from unstructured pathology reports using various language models, highlighting the challenges and potential of both proprietary and open-source solutions. A dataset comprising 579 annotated pathology reports in German and English was created to evaluate the performance of six language models, including the proprietary GPT-4 and several open-source models (Llama2 and Llama3 families, Qwen2.5). The study emphasizes the importance of prompt engineering and model quantization techniques for effective deployment on consumer-grade hardware, revealing that open-source models can achieve precision levels comparable to GPT-4.
The findings indicate that open-source language models not only provide a cost-effective and privacy-preserving alternative for structuring pathology data but also demonstrate significant variability in performance based on the configurations employed. The research underscores the potential for these models to facilitate the development of AI tools in healthcare, particularly in oncology, where vast amounts of valuable data are archived. The publicly available bilingual dataset serves as a benchmark for future research, aiming to enhance the systematic extraction and labeling of data necessary for training AI algorithms in the medical domain.
Methods
In this study, 579 pathology reports from 340 patients with prostate adenocarcinoma who underwent radical prostatectomy between 2020 and 2022 were analyzed to evaluate the performance of various language models in extracting structured data. The reports, retrieved in anonymized plain text format, were annotated based on established classifications, including the TNM Classification of Malignant Tumors and the Gleason grading system. Eleven clinically relevant parameters were selected for extraction, including WHO (ISUP) Grade Group, T-Stage, and lymph node involvement, which are critical for staging and treatment decisions. The reports were translated from German to English using the DeepL API, and the study adhered to ethical guidelines, with patient consent obtained for the use of anonymized data.
The evaluation involved comparing the structured data extraction capabilities of GPT-4 and Llama2/Llama3 models under various conditions. The models were tasked with extracting the specified parameters and providing outputs in structured JSON format. Performance metrics, including accuracy, precision, recall, and F1-Score, were calculated to assess the models’ effectiveness. The study also explored the feasibility of using quantized models for broader accessibility, allowing for structured data extraction on consumer-grade hardware. Overall, this research highlights the potential of language models in automating the extraction of critical information from pathology reports, which is essential for advancing AI applications in medical contexts.
Results
The evaluation of open-source and proprietary large language models (LLMs) demonstrated that Llama3, particularly the 70B variant, achieves accuracy comparable to GPT-4 in extracting structured data from pathology reports. Both models attained over 97% overall accuracy across eleven parameters in English and German. The smaller Llama3 8B model also performed well, achieving 91% and 83% accuracy in the respective languages. In contrast, the previous generation, Llama2, exhibited significantly lower performance regardless of model size. Accuracy analyses at the parameter level revealed that tumor staging and surgical outcomes were the most accurately extracted, while quantitative parameters posed more challenges, particularly for Llama2.
In terms of JSON output generation, GPT-4 produced valid JSON files for all reports, while Llama3 70B showed nearly flawless results. Llama2 struggled with JSON formatting, often yielding invalid outputs. Notably, a significant portion of invalid JSON files varied between initial and retest runs, indicating inherent issues with JSON generation rather than specific input problems. Additionally, processing times varied significantly, with GPT-4 averaging 12 seconds per report, while Llama3 8B was notably faster at 5-6 seconds. Llama2 70B, however, took considerably longer, averaging 51 seconds. Overall, the study highlights the superior performance and efficiency of Llama3 over its predecessor, Llama2, and positions both Llama3 and GPT-4 as effective tools for data extraction in pathology.
Discussion
In this study, the authors conducted a comprehensive evaluation of various large language models (LLMs) for the extraction of structured data from surgical pathology reports, specifically focusing on radical prostatectomy cases. They created a high-quality dataset with ground truth information and tested six different LLMs, including both proprietary (GPT-4) and open-source models (Llama2, Llama3, Qwen2.5). The results demonstrated that while GPT-4 achieved an impressive overall accuracy of 98% in English and 97% in German, open-source models like Llama3 also performed comparably well, achieving similar accuracy levels without the associated costs and privacy concerns of proprietary models.
The analysis highlighted the importance of prompt engineering and quantization techniques in enhancing model performance. Although quantized versions of LLMs (4-bit) showed a decline in accuracy compared to their full-weight counterparts, employing effective prompting strategies, particularly the chain-of-verification approach, improved results. The study found that hallucinations—instances of generating non-existent information—were more prevalent in Llama2 models, while GPT-4 exhibited minimal hallucinations. Additionally, the complexity of the text influenced extraction accuracy, with a correlation observed between lexical diversity and performance. Overall, the findings suggest that open-source LLMs can effectively structure pathology report data, offering a promising avenue for AI applications in the medical domain, albeit with room for improvement in quantized model performance.
