DOI: https://doi.org/10.1038/s41467-024-52417-z
PMID: https://pubmed.ncbi.nlm.nih.gov/39333468
تاريخ النشر: 2024-09-27
المؤلف: Pengcheng Qiu وآخرون
الموضوع الرئيسي: نمذجة الموضوعات
نظرة عامة
تقدم ورقة البحث تقدمًا كبيرًا في تطوير نماذج اللغة الطبية متعددة اللغات مفتوحة المصدر، والتي تهدف إلى خدمة جمهور لغوي متنوع. يقدم المؤلفون مجموعة طبية متعددة اللغات شاملة، تُعرف باسم MMedC، والتي تتكون من حوالي 25.5 مليار رمز عبر ست لغات رئيسية: الإنجليزية، الصينية، اليابانية، الفرنسية، الروسية، والإسبانية. تم تصميم هذه المجموعة لتسهيل التدريب التلقائي لنماذج اللغة الكبيرة (LLMs) في المجال الطبي. بالإضافة إلى ذلك، يقترح المؤلفون MMedBench، وهو معيار أسئلة متعددة الخيارات الطبية متعددة اللغات الذي يقيم قدرات النماذج المختلفة، مما يوفر إطارًا لمراقبة تقدم نماذج LLMs الطبية متعددة اللغات.
تظهر تقييمات عدة نماذج LLMs مفتوحة المصدر، بما في ذلك تلك التي تم تدريبها بشكل إضافي على MMedC، أن النموذج النهائي، MMed-Llama 3، الذي يحتوي على 8 مليارات معلمة فقط، يتفوق على جميع النماذج مفتوحة المصدر الأخرى في كل من MMedBench والمعايير الإنجليزية، حتى أنه يقترب من أداء النماذج المغلقة المصدر مثل GPT-4. تتناول هذه الدراسة القيود الحالية في أداء نماذج LLMs متعددة اللغات على الاستفسارات الطبية، لا سيما في اللغات غير الإنجليزية، من خلال توفير مجموعة بيانات قوية ومعيار يعكس التنوع اللغوي وتعقيد المجال الطبي. تؤكد النتائج على الإمكانية المتاحة لنماذج مفتوحة المصدر لسد الفجوة في الأداء مع نظرائها المغلقة المصدر، مما يعزز الوصول والفائدة في الرعاية الصحية عبر لغات مختلفة.
الطرق
في قسم الطرق، يوضح المؤلفون نتائجهم التجريبية والمنهجية المستخدمة في بحثهم. يظهرون أن التدريب التلقائي على مجموعة بيانات MMedC يعزز بشكل كبير أداء نماذج اللغة (LMs) مثل MMedLM، MMedLM 2، وMMed-Llama 3 مقارنة بنظيراتها الأساسية (InternLM، InternLM 2، وLlama 3). يبرز هذا التحسن القيود المفروضة على مجموعات البيانات السابقة الحالية في السياقات الطبية متعددة اللغات ويؤكد على ضرورة وجود مجموعات بيانات متخصصة مثل MMedC.
يؤكد المؤلفون أيضًا أن دمج بيانات متعددة اللغات عالية الجودة (HQ-Data) وبيانات مجموعة اللغة العامة المفلترة يساهم في تحقيق مكاسب في الأداء. تشير نتائجهم إلى أن دمج بيانات التبرير أثناء عملية الضبط الدقيق على MMedBench يؤدي إلى تحسينات ملحوظة في دقة المهام متعددة الخيارات، مع زيادات بنسبة 2.33% لـ InternLM، و2.42% لـ InternLM 2، و4.07% لـ Llama 3. بالإضافة إلى ذلك، يلاحظون أن نماذج LLMs الأساسية الأقوى، التي تم تدريبها مسبقًا على مجموعات بيانات أكبر وأكثر تنوعًا، تحقق نتائج أفضل في الإجابة على الأسئلة الطبية متعددة اللغات. يشير هذا إلى الحاجة إلى تطوير مجموعات بيانات طبية مفتوحة المصدر للاستفادة بشكل فعال من التقدم في LLMs العامة. يختتم القسم بمراجعة موجزة للمنهجية، بما في ذلك بناء MMedC، وإجراءات التدريب التلقائي، وتنظيم وتقييم المعيار الطبي متعدد اللغات الجديد، MMedBench.
النتائج
في هذا القسم، يقدم المؤلفون الخصائص الإحصائية لمجموعات البيانات التي قاموا بإنشائها لدراستهم. بعد ذلك، يقومون بتقييم نماذج اللغة الكبيرة (LLMs) المختلفة باستخدام معيار MMedBench، مع التركيز على أدائها في الأسئلة متعددة الخيارات وقدرات توليد التبرير. بالإضافة إلى ذلك، يتم تقييم فعالية مكون MMedC.
لفهم مساهمات كل مكون من مكونات مجموعة البيانات بشكل أفضل، يقوم المؤلفون بإجراء سلسلة من دراسات الإزالة. تهدف هذه الدراسات إلى توضيح تأثير العناصر الفردية داخل مجموعات البيانات على الأداء العام لنماذج LLMs، مما يوفر رؤى حول أهمية كل مكون في تعزيز قدرات النموذج.
المناقشة
في هذا القسم، يناقش المؤلفون تطوير وتقييم مجموعتين كبيرتين من البيانات: مجموعة البيانات الطبية متعددة اللغات (MMedC) ومعيار الأسئلة والأجوبة الطبية متعددة اللغات (MMedBench). تعتبر MMedC أكبر مجموعة بيانات طبية متعددة اللغات حتى الآن، حيث تتكون من أكثر من 25.5 مليار رمز مستمد من الكتب الطبية، والمواقع الإلكترونية، ومجموعات البيانات العامة المفلترة، تغطي ست لغات. يضمن تكوين مجموعة البيانات قابلية تطبيق واسعة عبر سياقات لغوية متنوعة، حيث تمثل الإنجليزية الحصة الأكبر بنسبة 42%. يعمل MMedBench كمعيار لتقييم نماذج الطب متعددة اللغات، ويتميز بمجموعة متنوعة من الأسئلة الطبية المصنفة إلى 21 موضوعًا، ويشمل أسئلة متعددة الخيارات مع أقسام تبرير لتعزيز تدريب النموذج وتقييمه.
يكشف تقييم نماذج اللغة الكبيرة (LLMs) المختلفة على MMedBench أنه بينما تؤدي النماذج الاحتكارية مثل GPT-4 بشكل جيد، كانت النماذج مفتوحة المصدر متأخرة في البداية لكنها تحسنت بشكل كبير بعد الضبط الدقيق على MMedC. من الجدير بالذكر أن نموذج MMed-Llama 3، الذي تم ضبطه بدقة على MMedC، حقق درجات دقة تنافسية، لا سيما في الإنجليزية، وأظهر قدرات محسنة في توليد التبرير. تؤكد الدراسة أيضًا على أهمية نماذج LLMs متعددة اللغات في معالجة حواجز اللغة في الرعاية الصحية، وتحسين التعليم الطبي، وتسهيل التواصل الأفضل بين المرضى ومقدمي الرعاية. تسلط النتائج الضوء على الإمكانية المتاحة لهذه النماذج للمساهمة في تطوير الذكاء الاصطناعي الطبي العام (GMAI) وتعزيز طرق التوليد المعززة بالاسترجاع، مما يحسن في النهاية تقديم الرعاية الصحية عبر الحدود اللغوية والثقافية.
القيود
في هذا القسم، يعترف المؤلفون بعدة قيود واجهوها خلال بحثهم حول بناء مجموعة بيانات طبية متعددة اللغات وتعزيز قدرات نماذج اللغة (LLMs) للتطبيقات الطبية. أولاً، يثير الاعتماد على بيانات مستخرجة من الويب مخاوف بشأن التحيزات الكامنة، لا سيما ضد الفئات المحرومة، مما يشكل تحديًا كبيرًا في تطوير نماذج LLMs الطبية العادلة. يعبر المؤلفون عن نيتهم في تنفيذ ضوابط أمان أكثر صرامة لمعالجة هذه التحيزات في الأعمال المستقبلية.
ثانيًا، بينما تم بذل جهود لتحسين قابلية تفسير النموذج، يشير المؤلفون إلى أن تطوير آليات قابلية التفسير لهياكل LLM لا يزال غير مستكشف بشكل كافٍ. بالإضافة إلى ذلك، لا تشمل مجموعة البيانات الحالية جميع اللغات المنطوقة عالميًا، مما يدفع إلى خطط لتوسيع تغطية اللغات، بما في ذلك الألمانية والعربية، من خلال الاستفادة من مجموعات بيانات الزحف الشائعة. كما يبرز المؤلفون إمكانية دمج الأدبيات الطبية المحلية لإثراء مجموعة التدريب بشكل أكبر. أخيرًا، يذكرون القيود الحاسوبية لنموذجهم الحالي الذي يحتوي على 8 مليارات معلمة ويقترحون الانتقال إلى هيكل أكبر مع تعزيز الاسترجاع في التكرارات المستقبلية لتحسين الأداء والتخفيف من المشكلات المتعلقة بالهلوسة.
DOI: https://doi.org/10.1038/s41467-024-52417-z
PMID: https://pubmed.ncbi.nlm.nih.gov/39333468
Publication Date: 2024-09-27
Author(s): Pengcheng Qiu et al.
Primary Topic: Topic Modeling
Overview
The research paper presents significant advancements in the development of open-source, multilingual medical language models aimed at serving a diverse linguistic audience. The authors introduce a comprehensive multilingual medical corpus, referred to as MMedC, which consists of approximately 25.5 billion tokens across six primary languages: English, Chinese, Japanese, French, Russian, and Spanish. This corpus is designed to facilitate autoregressive training for large language models (LLMs) in the medical domain. Additionally, the authors propose MMedBench, a multilingual medical multi-choice question-answering benchmark that evaluates the rationale capabilities of various models, thereby providing a framework for monitoring the progress of multilingual medical LLMs.
The evaluation of several open-source LLMs, including those further trained on MMedC, reveals that the final model, MMed-Llama 3, with only 8 billion parameters, outperforms all other open-source models on both MMedBench and English benchmarks, even approaching the performance of closed-source models like GPT-4. This work addresses the existing limitations in the performance of multilingual LLMs on medical queries, particularly in non-English languages, by providing a robust dataset and benchmark that reflect the linguistic diversity and complexity of the medical field. The findings underscore the potential for open-source models to bridge the performance gap with their closed-source counterparts, thereby enhancing accessibility and utility in healthcare across different languages.
Methods
In the Methods section, the authors outline their experimental findings and the methodology employed in their research. They demonstrate that auto-regressive training on the MMedC dataset significantly enhances the performance of language models (LMs) such as MMedLM, MMedLM 2, and MMed-Llama 3 compared to their baseline counterparts (InternLM, InternLM 2, and Llama 3). This improvement highlights the limitations of existing pre-training corpora in multilingual medical contexts and emphasizes the necessity for specialized datasets like MMedC.
The authors further establish that incorporating high-quality multilingual data (HQ-Data) and filtered general language corpus data contributes to performance gains. Their results indicate that integrating rationale data during the fine-tuning process on MMedBench leads to notable accuracy improvements in multiple-choice tasks, with increases of 2.33% for InternLM, 2.42% for InternLM 2, and 4.07% for Llama 3. Additionally, they observe that stronger foundational LLMs, which have been pretrained on larger and more diverse corpora, yield better results in multilingual medical question answering. This suggests a need for the development of open-source medical datasets to leverage advancements in general LLMs effectively. The section concludes with a brief overview of the methodology, including the construction of MMedC, the auto-regressive training procedure, and the curation and evaluation of the new multilingual medical benchmark, MMedBench.
Results
In this section, the authors present the statistical characteristics of the datasets they constructed for their study. They subsequently evaluate various large language models (LLMs) using the MMedBench benchmark, focusing on their performance in multi-choice questions and rationale generation capabilities. Additionally, the effectiveness of the MMedC component is assessed.
To further understand the contributions of each dataset component, the authors conduct a series of ablation studies. These studies aim to elucidate the impact of individual elements within the datasets on the overall performance of the LLMs, thereby providing insights into the significance of each component in enhancing model capabilities.
Discussion
In this section, the authors discuss the development and evaluation of two significant datasets: the Multilingual Medical Corpus (MMedC) and the Multilingual Medical Question and Answering Benchmark (MMedBench). MMedC is the largest multilingual medical corpus to date, comprising over 25.5 billion tokens sourced from medical textbooks, websites, and filtered general corpora, covering six languages. The dataset’s composition ensures broad applicability across diverse linguistic contexts, with English representing the largest share at 42%. MMedBench serves as a benchmark for evaluating multilingual medical models, featuring a variety of medical questions categorized into 21 topics, and includes multiple-choice questions with rationale sections to enhance model training and evaluation.
The evaluation of various large language models (LLMs) on MMedBench reveals that while proprietary models like GPT-4 perform well, open-source models initially lagged but improved significantly after fine-tuning on MMedC. Notably, the MMed-Llama 3 model, fine-tuned on MMedC, achieved competitive accuracy scores, particularly in English, and demonstrated enhanced rationale generation capabilities. The study also emphasizes the importance of multilingual LLMs in addressing language barriers in healthcare, improving medical education, and facilitating better communication between patients and providers. The findings highlight the potential for these models to contribute to the development of General Medical Artificial Intelligence (GMAI) and enhance retrieval-augmented generation methods, ultimately improving healthcare delivery across linguistic and cultural boundaries.
Limitations
In this section, the authors acknowledge several limitations encountered during their research on constructing a multilingual medical corpus and enhancing the capabilities of language models (LLMs) for medical applications. Firstly, the reliance on web-crawled data raises concerns about inherent biases, particularly against underprivileged populations, which poses a significant challenge in developing equitable medical LLMs. The authors express their intention to implement more stringent safety controls to address these biases in future work.
Secondly, while efforts have been made to improve the explainability of the model, the authors note that the development of explainability mechanisms for LLM architectures remains under-explored. Additionally, the current dataset does not encompass all languages spoken globally, prompting plans to expand language coverage, including German and Arabic, by leveraging common crawl datasets. The authors also highlight the potential of integrating local medical literature to enrich the training corpus further. Lastly, they mention the computational constraints of their current 8B scale model and propose transitioning to a larger architecture with retrieval augmentation in future iterations to enhance performance and mitigate issues related to hallucinations.
