DOI: https://doi.org/10.1038/s41746-024-01390-4
PMID: https://pubmed.ncbi.nlm.nih.gov/39865143
تاريخ النشر: 2025-01-27
المؤلف: Chaoyi Wu وآخرون
الموضوع الرئيسي: نمذجة الموضوعات
نظرة عامة
في هذه الدراسة، يقدم المؤلفون MedS-Bench، وهو معيار شامل مصمم لتقييم نماذج اللغة الكبيرة (LLMs) في السياقات السريرية عبر 11 مهمة معقدة، بما في ذلك تلخيص التقارير السريرية، وتوصيات العلاج، والتشخيص، والتعرف على الكيانات المسماة. يكشف تقييم تسعة من نماذج LLM الرائدة، مثل MEDITRON وLlama 3 وGPT-4، أن هذه النماذج غالبًا ما تواجه صعوبة في تعقيدات المهام السريرية، على الرغم من أدائها العالي في معايير الإجابة على الأسئلة متعددة الخيارات (MCQA). تسلط هذه الفجوة الضوء على عدم كفاية المعايير الحالية في عكس الفائدة العملية لنماذج LLM في السيناريوهات السريرية الواقعية.
لمعالجة هذه القيود، طور المؤلفون MedS-Ins، وهو مجموعة بيانات كبيرة لتوجيه التعليمات تتكون من 58 مجموعة نصوص لغوية موجهة طبيًا تحتوي على 5 ملايين حالة و19,000 تعليمات عبر 122 مهمة. أظهر تجربة إثبات المفهوم فائدة مجموعة البيانات من خلال توجيه تعليمات نموذج لغة طبية خفيف الوزن، مما أدى إلى MMedIns-Llama 3، الذي تفوق بشكل كبير على النماذج الموجودة في مختلف المهام السريرية. يؤكد المؤلفون على أهمية جعل MedS-Ins متاحًا لمجتمع البحث لتعزيز المزيد من التقدم في تطوير نماذج LLM الطبية مفتوحة المصدر القادرة على معالجة مجموعة واسعة من التحديات السريرية. بالإضافة إلى ذلك، أطلقوا لوحة متقدمة ديناميكية لـ MedS-Bench لتتبع التقدم في هذا المجال.
الطرق
في هذا القسم، يوضح المؤلفون المنهجية المستخدمة لبناء مجموعة بيانات MedS-Ins، كما هو موضح في الشكل 3a. تتضمن العملية تصنيف المهام من خلال تعيين علامة مجال، تشير إلى المجال الطبي المعني، وعلامة فئة، تحدد المهمة القابلة للتطبيق. في البداية، يقوم الباحثون بتصفية الجمل المتعلقة بالتعليمات الطبية من مجموعات بيانات التعليمات الطبيعية الموجودة. بعد ذلك، يستخدمون تقنيات BioNLP المستهدفة لتحويل هذه الجمل المصفاة إلى تنسيقات استجابة نصية حرة، مما يعزز فائدة مجموعة البيانات لمختلف التطبيقات الطبية.
النتائج
في هذا القسم، يقدم المؤلفون نتائج دراستهم، مع التركيز على تطوير وتقييم MedS-Bench، وهو معيار مصمم للتقييم الشامل للتطبيقات السريرية في معالجة اللغة الطبية. تشمل MedS-Bench 11 فئة مهمة سريرية عالية المستوى مستمدة من 28 مجموعة بيانات، أعيد تنسيقها إلى هيكل سؤال-جواب مدفوع بالتعليمات. يقدم المؤلفون أيضًا MedS-Ins، وهي مجموعة بيانات لتوجيه التعليمات تتكون من 122 مهمة سريرية متميزة عبر خمسة مجالات نصية: الامتحانات، والنصوص السريرية، والأوراق الأكاديمية، وقواعد المعرفة الطبية، والمحادثات اليومية. يساهم كل مجال بنسب متفاوتة من الرموز، حيث تشكل النصوص السريرية 35% وقواعد المعرفة الطبية 43%، مما يضمن التوافق مع المتطلبات السريرية.
تقوم التحليل أيضًا بتصنيف المهام بناءً على المهارات المطلوبة، مميزًا بين استرجاع الحقائق من النموذج واسترجاع الحقائق من السياق المقدم. يحدد المؤلفون 19 فئة مهمة أساسية لنماذج اللغة الطبية (LLMs)، والتي تشمل ليس فقط الفئات الـ 11 في MedS-Bench ولكن أيضًا مهام لغوية وتحليلية إضافية. يتم تقييم أداء نماذج LLM المختلفة، بما في ذلك نموذج المؤلفين MMedIns-Llama 3، باستخدام استراتيجية تحفيز من 3 لقطات لمعظم المهام، بينما يتم استخدام نهج عدم اللقطات لمهام الإجابة على الأسئلة متعددة الخيارات (MCQA). تشير النتائج إلى مقارنة شاملة للنماذج الرائدة، بما في ذلك النماذج الملكية مثل GPT-4، مع التركيز على إدارة قيود التكلفة من خلال أخذ عينات عشوائية من حالات الاختبار.
المناقشة
في هذا القسم، يناقش المؤلفون تقييم نموذجهم المقترح، MMedIns-Llama 3، عبر مختلف المهام السريرية، مع تسليط الضوء على أدائه مقارنة بنماذج اللغة الكبيرة (LLMs) الموجودة. تشير النتائج إلى أنه بينما تتفوق النماذج الملكية مثل GPT-4 وClaude-3.5 في معايير الإجابة على الأسئلة متعددة الخيارات (MCQA)، إلا أنها تظهر قيودًا في المهام السريرية الأكثر تعقيدًا مثل التشخيص، وتخطيط العلاج، وتوقع النتائج السريرية. يظهر MMedIns-Llama 3، على الرغم من عدم تدريبه بشكل أساسي على MCQA، تحسينات كبيرة في الدقة عبر معايير متنوعة، محققًا درجات ملحوظة في استخراج المعلومات، وتلخيص النصوص، والتعرف على الكيانات المسماة (NER)، متفوقًا على العديد من النماذج الموجودة، بما في ذلك GPT-4 في مهام معينة.
يؤكد المؤلفون على ضرورة وجود إطار تقييم شامل، MedS-Bench، الذي يقدمونه لتقييم LLMs في مجموعة أوسع من السيناريوهات السريرية بخلاف MCQA. كما يقدمون MedS-Ins، وهي مجموعة بيانات جديدة لتوجيه التعليمات مصممة لتعزيز توافق LLM مع المتطلبات السريرية من خلال دمج مجموعة متنوعة من النصوص الطبية. على الرغم من النتائج الواعدة، يعترف المؤلفون بالقيود في معاييرهم ومجموعات البيانات الحالية، مقترحين العمل المستقبلي لتوسيع نطاق المهام واللغات، وإجراء تحقق سريري لضمان التطبيق العملي لنموذجهم في الإعدادات الواقعية. بشكل عام، تؤكد النتائج على إمكانيات MMedIns-Llama 3 كأداة قوية لمجموعة متنوعة من التطبيقات الطبية، بينما تدعو أيضًا إلى تطوير مستمر وتعاون مجتمعي لتحسين نماذج LLM الطبية بشكل أكبر.
DOI: https://doi.org/10.1038/s41746-024-01390-4
PMID: https://pubmed.ncbi.nlm.nih.gov/39865143
Publication Date: 2025-01-27
Author(s): Chaoyi Wu et al.
Primary Topic: Topic Modeling
Overview
In this study, the authors introduce MedS-Bench, a comprehensive benchmark designed to evaluate large language models (LLMs) in clinical contexts across 11 complex tasks, including clinical report summarization, treatment recommendations, diagnosis, and named entity recognition. The evaluation of nine leading LLMs, such as MEDITRON, Llama 3, and GPT-4, reveals that these models often struggle with the intricacies of clinical tasks, despite their high performance on multiple-choice question-answering (MCQA) benchmarks. This discrepancy highlights the inadequacy of current benchmarks in reflecting the practical utility of LLMs in real-world clinical scenarios.
To address these limitations, the authors developed MedS-Ins, a large-scale instruction-tuning dataset comprising 58 medically oriented language corpora with 5 million instances and 19,000 instructions across 122 tasks. A proof-of-concept experiment demonstrated the dataset’s utility by instruction tuning a lightweight medical language model, resulting in MMedIns-Llama 3, which significantly outperformed existing models on various clinical tasks. The authors emphasize the importance of making MedS-Ins accessible to the research community to foster further advancements in the development of open-source medical LLMs capable of addressing a wide range of clinical challenges. Additionally, they launched a dynamic leaderboard for MedS-Bench to track progress in the field.
Methods
In this section, the authors outline the methodology employed for constructing the MedS-Ins dataset, as illustrated in Figure 3a. The process involves categorizing tasks by assigning a domain tag, which indicates the relevant medical domain, and a category tag, which specifies the applicable task. Initially, the researchers filter sentences related to medical instructions from existing natural instruction datasets. Subsequently, they utilize targeted BioNLP techniques to convert these filtered sentences into free-text response formats, thereby enhancing the dataset’s utility for various medical applications.
Results
In this section, the authors present the results of their study, focusing on the development and evaluation of MedS-Bench, a benchmark designed for comprehensive assessment of clinical applications in medical language processing. MedS-Bench includes 11 high-level clinical task categories derived from 28 datasets, reformatted into an instruction-prompted question-answering structure. The authors also introduce MedS-Ins, an instruction tuning dataset comprising 122 distinct clinical tasks across five text domains: exams, clinical texts, academic papers, medical knowledge bases, and daily conversations. Each domain contributes varying proportions of tokens, with clinical texts making up 35% and medical knowledge bases 43%, ensuring alignment with clinical demands.
The analysis further categorizes tasks based on the skills required, distinguishing between recalling facts from the model and retrieving facts from provided context. The authors identify 19 task categories essential for medical language models (LLMs), which include not only the 11 categories in MedS-Bench but also additional linguistic and analytical tasks. The performance of various LLMs, including the authors’ model MMedIns-Llama 3, is evaluated using a 3-shot prompting strategy for most tasks, while a zero-shot approach is used for multiple-choice question answering (MCQA) tasks. The results indicate a thorough comparison of leading models, including proprietary ones like GPT-4, with a focus on managing cost constraints through a random sampling of test cases.
Discussion
In this section, the authors discuss the evaluation of their proposed model, MMedIns-Llama 3, across various clinical tasks, highlighting its performance compared to existing large language models (LLMs). The results indicate that while proprietary models like GPT-4 and Claude-3.5 excel in multiple-choice question-answering (MCQA) benchmarks, they exhibit limitations in more complex clinical tasks such as diagnosis, treatment planning, and clinical outcome prediction. MMedIns-Llama 3, despite not being primarily trained on MCQA, demonstrates significant improvements in accuracy across diverse benchmarks, achieving notable scores in information extraction, text summarization, and named entity recognition (NER), outperforming many existing models, including GPT-4 in specific tasks.
The authors emphasize the necessity of a comprehensive evaluation framework, MedS-Bench, which they introduce to assess LLMs in a broader range of clinical scenarios beyond MCQA. They also present MedS-Ins, a novel instruction tuning dataset designed to enhance LLM alignment with clinical demands by incorporating a diverse array of medical texts. Despite the promising results, the authors acknowledge limitations in their current benchmarks and datasets, suggesting future work to expand the range of tasks and languages, and to conduct clinical validations to ensure the practical applicability of their model in real-world settings. Overall, the findings underscore the potential of MMedIns-Llama 3 as a robust tool for various medical applications, while also calling for ongoing development and community collaboration to refine medical LLMs further.
