DOI: https://doi.org/10.1038/s41598-024-64827-6
PMID: https://pubmed.ncbi.nlm.nih.gov/38898116
تاريخ النشر: 2024-06-19
المؤلف: Jenish Maharjan وآخرون
الموضوع الرئيسي: نمذجة الموضوعات
نظرة عامة
تقدم البحث OpenMedLM، وهي منصة مفتوحة المصدر مصممة لتعزيز أداء نماذج اللغة الكبيرة (LLMs) في التطبيقات الطبية. بينما تتطلب النماذج التقليدية المملوكة تعديلات دقيقة واسعة وبيانات متخصصة، يظهر OpenMedLM أن نماذج الأساس مفتوحة المصدر يمكن أن تحقق نتائج رائدة (SOTA) دون هذه القيود. قامت الدراسة بتقييم مجموعة متنوعة من نماذج LLMs مفتوحة المصدر تتراوح من 7B إلى 70B معلمات على معايير طبية، بما في ذلك MedQA وMMLU مجموعة فرعية طبية، واختيار نموذج Yi34B في النهاية. حقق OpenMedLM دقة ملحوظة بلغت 72.6% على MedQA و81.7% على MMLU، متجاوزًا المعايير السابقة SOTA وعرض إمكانيات النماذج مفتوحة المصدر في الرعاية الصحية.
تؤكد النتائج على الخصائص الناشئة لنماذج LLMs مفتوحة المصدر في السياقات الطبية وتدعو لاستخدام استراتيجيات تحفيز مبتكرة، مثل التعلم بدون أمثلة والتعلم من أمثلة قليلة، لتعزيز أداء النموذج بشكل أكبر. يجادل المؤلفون بأن OpenMedLM لا يوفر فقط طريقًا لتحسين الوصول وقابلية التخصيص في تطبيقات الرعاية الصحية، بل يشجع أيضًا على البحث المستمر في تقنيات هندسة التحفيز. من خلال دمج قدرات LLM مع خوارزميات الذكاء الاصطناعي الطبية الموجودة، تقترح الدراسة أن هذه النماذج يمكن تنفيذها بشكل فعال في البيئات السريرية، مما يساهم في تحسين نتائج الصحة للمرضى.
الطرق
في هذا القسم، يوضح المؤلفون طرق التقييم الخاصة بهم لتقييم أداء نموذجهم على معايير الأسئلة والأجوبة (Q&A) متعددة الخيارات، باستخدام الدقة كمقياس رئيسي. تم السماح للنموذج بخمس محاولات كحد أقصى لتوليد مخرجات صالحة، تعرف على أنها صيغة إجابة يمكن التعرف عليها كاستجابة متعددة الخيارات (مثل، “(A)”). إذا فشل النموذج في إنتاج مخرج صالح ضمن هذه المحاولات، تم اعتبار الاستجابة غير صحيحة، مما يؤثر على درجة الدقة العامة. تم تحديد الحد الأقصى لعدد الرموز لمخرجات النموذج تجريبيًا لتحقيق توازن بين متوسط طول تفسيرات سلسلة التفكير (CoT) مع استخدام ذاكرة GPU ووقت التقييم.
لتحسين الأداء، تم اختبار تعليمات تحفيز متنوعة، بما في ذلك تحفيزات طبية مخصصة وصيغ Q&A المعتمدة من الأدبيات السابقة. قام المؤلفون بتنفيذ طريقة الجيران الأقرب (kNN) لتحديد خمسة أسئلة تدريبية مشابهة لكل سؤال اختبار، مما يعزز الصلة السياقية. بالإضافة إلى ذلك، تم استخدام مخطط تصويت جماعي، حيث تمت معالجة كل تحفيز خمس مرات مع خيارات إجابة عشوائية، وتم اختيار الإجابة الأكثر تكرارًا. تم توليد مخرجات النموذج بدرجة حرارة فك تشفير تبلغ 0.4، مما يعزز التوازن بين الاستجابات الحتمية والمتنوعة. تم إجراء دراسة إلغاء لتقييم فعالية كل تقنية تحفيز بشكل فردي قبل تقييم الأداء الشامل لمنصة تحفيز OpenMedLM على نموذج الأساس OS Yi 34B.
النتائج
تظهر النتائج أن نموذج الأساس Yi 34B، باستخدام منصة تحفيز OpenMedLM الكاملة، يحقق أداءً رائدًا (SOTA) في ثلاثة من أربعة معايير للأسئلة والأجوبة الطبية. كما هو موضح في الجدول 1، تظهر دقة نموذج Yi 34B عبر معايير مختلفة، بما في ذلك MedQA وMedMCQA وPubMedQA، تحسنًا كبيرًا عند استخدام مجموعة كاملة من استراتيجيات التحفيز. على وجه التحديد، يصل نموذج Yi 34B إلى دقة تبلغ 72.6% على معيار MMLU-Medical، متجاوزًا أداء نموذج Meditron 70B في هذا السياق.
توضح دراسة الإلغاء بشكل أكبر مساهمات المكونات الفردية داخل منصة تحفيز OpenMedLM، كاشفة أنه بينما يحقق نموذج Meditron دقة أعلى في سيناريوهات بدون أمثلة، فإن استراتيجيات التحفيز الشاملة لـ OpenMedLM تسمح لنموذج Yi 34B بالتفوق بشكل عام. توضح الشكل 2 الفعالية المقارنة لكل مكون فرعي من OpenMedLM مقابل تقنيات التحفيز المستخدمة من قبل Meditron، مما يبرز القدرات المعززة لنموذج Yi 34B عندما يتم دمج جميع المكونات.
المناقشة
يقدم البحث OpenMedLM، وهي منصة تحفيز تعزز أداء نماذج اللغة الأساسية مفتوحة المصدر (OS) على المعايير الطبية، محققة نتائج رائدة (SOTA). أظهر نموذج Yi 34B، باستخدام OpenMedLM، تحسينات كبيرة في الدقة عبر مجموعات بيانات مختلفة، بما في ذلك MedQA وMedMCQA وMMLU مجموعة فرعية طبية. كانت الأداء الأولي بدون أمثلة 58.4%، والذي تحسن إلى 72.6% مع مجموعة من استراتيجيات التحفيز، بما في ذلك التحفيز العشوائي من أمثلة قليلة، وتفسيرات سلسلة التفكير (CoT)، وتقنيات الجيران الأقرب (kNN). من الجدير بالذكر أن الجمع بين CoT مع التحفيز من أمثلة قليلة باستخدام kNN حقق أعلى دقة، مما يبرز الفوائد التراكمية والتآزرية لهذه التقنيات.
تشير النتائج إلى أن النماذج مفتوحة المصدر يمكن أن تحقق أداءً تنافسيًا مقارنة بالنماذج المتخصصة، مثل Meditron، دون الحاجة إلى تعديلات دقيقة واسعة. لم تحسن استراتيجيات التحفيز لـ OpenMedLM أداء نموذج Yi 34B فحسب، بل سلطت الضوء أيضًا على الإمكانيات للخصائص الناشئة في النماذج مفتوحة المصدر، مما يشير إلى أن هندسة التحفيز القوية يمكن أن تفتح إمكانيات كانت تُعتقد سابقًا أنها تتطلب تعديلات دقيقة. تدعو النتائج إلى مزيد من الاستكشاف للنماذج مفتوحة المصدر في تطبيقات الرعاية الصحية، مما يبرز إمكانية الوصول والتكيف، وهي أمور حاسمة للتنفيذ في العالم الحقيقي في البيئات السريرية. بشكل عام، تضع الدراسة الأساس لاستغلال النماذج مفتوحة المصدر في المهام المعقدة للرعاية الصحية، مما يعزز التحول نحو حلول ذكاء اصطناعي أكثر شمولية وشفافية في الطب.
القيود
تسلط القيود الموضحة في هذا القسم الضوء على التحديات التي تواجه نماذج اللغة الكبيرة مفتوحة المصدر (OS) الرائدة (SOTA) في التطبيقات العملية. بينما تظهر النماذج أداءً قويًا في معايير الأسئلة والأجوبة الأكاديمية، لا تمثل هذه التقييمات بشكل كامل الاستفسارات المتنوعة والدقيقة التي تواجهها في البيئات السريرية. على وجه الخصوص، يحد الاعتماد على الأسئلة متعددة الخيارات من قدرة النماذج على الانخراط في حوار مفتوح، وهو أمر غالبًا ما يكون ضروريًا لمعالجة تعقيدات التفاعلات الصحية في العالم الحقيقي.
علاوة على ذلك، تزداد إمكانية حدوث الهلاوس، أو الاستجابات غير المكتملة، أو التحيزات مع تكليف النماذج بتوليد مخرجات أكثر تعقيدًا تتطلب التفكير والأدلة. تؤكد هذه التعقيدات على الحاجة إلى مزيد من البحث والتجريب لتعزيز قدرات النماذج. يبقى المؤلفون متفائلين بأنه مع استمرار التطوير، ستتطور نماذج LLMs مفتوحة المصدر للتعامل بفعالية مع الاستفسارات التي تركز على الأطباء وتلك التي تركز على المرضى، مما يسهل في النهاية دمجها في سيناريوهات الرعاية الصحية في العالم الحقيقي.
DOI: https://doi.org/10.1038/s41598-024-64827-6
PMID: https://pubmed.ncbi.nlm.nih.gov/38898116
Publication Date: 2024-06-19
Author(s): Jenish Maharjan et al.
Primary Topic: Topic Modeling
Overview
The research presents OpenMedLM, an open-source prompting platform designed to enhance the performance of large language models (LLMs) in medical applications. While traditional proprietary models require extensive fine-tuning and specialized data, OpenMedLM demonstrates that open-source foundation models can achieve state-of-the-art (SOTA) results without such constraints. The study evaluated various open-source LLMs ranging from 7B to 70B parameters on medical benchmarks, including MedQA and MMLU medical-subset, ultimately selecting the Yi34B model. OpenMedLM achieved notable accuracies of 72.6% on MedQA and 81.7% on MMLU, surpassing previous SOTA benchmarks and showcasing the potential of open-source models in healthcare.
The findings emphasize the emergent properties of open-source LLMs in medical contexts and advocate for the use of innovative prompting strategies, such as zero-shot and few-shot learning, to further enhance model performance. The authors argue that OpenMedLM not only provides a pathway for improved accessibility and customizability in healthcare applications but also encourages ongoing research into prompt engineering techniques. By integrating LLM capabilities with existing medical AI algorithms, the study suggests that these models can be effectively implemented in clinical settings, ultimately contributing to better health outcomes for patients.
Methods
In this section, the authors detail their evaluation methods for assessing the performance of their model on multiple choice question and answer (Q&A) benchmarks, using accuracy as the primary metric. The model was allowed a maximum of five attempts to generate a valid output, defined as an answer format recognizable as a multiple choice response (e.g., “(A)”). If the model failed to produce a valid output within these attempts, the response was marked incorrect, impacting the overall accuracy score. The maximum number of tokens for model output was empirically determined to balance the average length of chain-of-thought (CoT) explanations with GPU memory usage and evaluation time.
To optimize performance, various prompting instructions were tested, including customized medical prompts and established Q&A formats from prior literature. The authors implemented a k-nearest neighbors (kNN) method to identify the five most similar training questions for each test question, enhancing contextual relevance. Additionally, an ensemble voting scheme was employed, where each prompt was processed five times with randomized answer options, and the most frequently output answer was selected. The model’s outputs were generated with a decoding temperature of 0.4, promoting a balance between deterministic and diverse responses. An ablation study was conducted to evaluate the effectiveness of each prompting technique individually before assessing the comprehensive performance of the OpenMedLM prompting platform on the OS Yi 34B foundation model.
Results
The results demonstrate that the Yi 34B foundation model, utilizing the complete OpenMedLM prompting platform, achieves state-of-the-art (SOTA) performance on three out of four medical question-and-answer benchmarks. As detailed in Table 1, the Yi 34B model’s accuracy across various benchmarks, including MedQA, MedMCQA, and PubMedQA, shows significant improvement when employing the full suite of prompting strategies. Specifically, the Yi 34B model reaches an accuracy of 72.6% on the MMLU-Medical benchmark, surpassing the Meditron 70B model’s performance in this context.
The ablation study further elucidates the contributions of individual components within the OpenMedLM prompting platform, revealing that while the Meditron model achieves higher accuracy in zero-shot scenarios, the comprehensive prompting strategies of OpenMedLM allow the Yi 34B model to excel overall. Figure 2 illustrates the comparative effectiveness of each sub-component of OpenMedLM against the prompting techniques used by Meditron, underscoring the enhanced capabilities of the Yi 34B model when all components are integrated.
Discussion
The research presents OpenMedLM, a prompting platform that enhances the performance of open-source (OS) foundation language models (LLMs) on medical benchmarks, achieving state-of-the-art (SOTA) results. The Yi 34B model, utilizing OpenMedLM, demonstrated significant improvements in accuracy across various datasets, including MedQA, MedMCQA, and the medical subset of MMLU. Initial zero-shot performance was 58.4%, which improved to 72.6% with a combination of prompting strategies, including random few-shot prompting, chain-of-thought (CoT) explanations, and k-nearest neighbors (kNN) techniques. Notably, the combination of CoT with kNN few-shot prompting yielded the highest accuracy, underscoring the additive and synergistic benefits of these techniques.
The findings indicate that OS models can achieve competitive performance compared to specialized models, such as Meditron, without the need for extensive fine-tuning. OpenMedLM’s prompting strategies not only improved the Yi 34B model’s performance but also highlighted the potential for emergent properties in OS models, suggesting that robust prompt engineering can unlock capabilities previously thought to require fine-tuning. The results advocate for further exploration of OS models in healthcare applications, emphasizing their accessibility and adaptability, which are crucial for real-world implementation in clinical settings. Overall, the study lays the groundwork for leveraging OS models in complex healthcare tasks, promoting a shift towards more inclusive and transparent AI solutions in medicine.
Limitations
The limitations outlined in this section highlight the challenges faced by the state-of-the-art (SOTA) open-source (OS) foundation large language models (LLMs) in practical applications. While the models demonstrate strong performance in academic question-and-answer benchmarks, these evaluations do not fully represent the diverse and nuanced inquiries encountered in clinical settings. Specifically, the reliance on multiple-choice questions restricts the models’ ability to engage in open-ended dialogue, which is often necessary for addressing the complexities of real-world healthcare interactions.
Moreover, the potential for hallucinations, incomplete responses, or biases increases as the models are tasked with generating more sophisticated outputs that require reasoning and evidence. This complexity underscores the need for further research and experimentation to enhance the models’ capabilities. The authors remain optimistic that with continued development, OS LLMs will evolve to effectively handle both physician-centric and patient-centric inquiries, ultimately facilitating their integration into real-world healthcare scenarios.
