DOI: https://doi.org/10.3389/fpubh.2025.1635381
PMID: https://pubmed.ncbi.nlm.nih.gov/41132171
تاريخ النشر: 2025-10-08
المؤلف: Shan Xu وآخرون
الموضوع الرئيسي: نمذجة الموضوعات
نظرة عامة
تقدم ورقة البحث إطارًا جديدًا، يسمى MEGA-RAG (تحسين الإجابات الموجهة بالعديد من الأدلة)، يهدف إلى معالجة مشكلة الهلوسات في نماذج اللغة الكبيرة (LLMs) المستخدمة في الصحة العامة. تشير الهلوسات إلى المخرجات غير الدقيقة من الناحية الواقعية التي يمكن أن تقوض التواصل السريري وصنع السياسات. يدمج MEGA-RAG تقنيات استرجاع الأدلة من مصادر متعددة، بما في ذلك الاسترجاع الكثيف عبر FAISS، والاسترجاع القائم على الكلمات الرئيسية عبر BM، ورسوم بيانية للمعرفة الطبية الحيوية، إلى جانب إعادة ترتيب عبر مشفر متقاطع للملاءمة الدلالية ووحدة تحسين مدركة للاختلاف لتعزيز الدقة الواقعية. تشير النتائج التجريبية إلى أن MEGA-RAG يقلل بشكل كبير من معدلات الهلوسة ويتفوق على النماذج الأساسية مثل PubMedBERT وPubMedGPT عبر مقاييس رئيسية، بما في ذلك الدقة، والموثوقية، والاسترجاع، ودرجة F1.
تصميم الإطار المعياري لا يحسن الأداء فحسب، بل يعزز أيضًا الشفافية، مما يسمح للمستخدمين بتتبع أصول كل ادعاء قدمه النموذج. يتم وضع MEGA-RAG كأداة قيمة لمجموعة متنوعة من تطبيقات الصحة العامة، بما في ذلك مراقبة تفشي الأمراض ودعم اتخاذ القرار السريري، من خلال تقديم استجابات قائمة على الأدلة وذات صلة بالسياق. تشمل اتجاهات البحث المستقبلية توسيع التقييمات لتشمل معايير أكبر، ودمج التحديثات الديناميكية من الوكالات الصحية، ودعم النشر متعدد اللغات، وتعزيز وحدة التوضيح الذاتي بتعليقات الأطباء. بالإضافة إلى ذلك، يخطط المؤلفون لاستكشاف استراتيجيات نشر فعالة مناسبة للبيئات ذات الموارد المحدودة، مما يضمن بقاء الإطار فعالًا في إعدادات متنوعة.
مقدمة
تسلط المقدمة الضوء على الزيادة المتزايدة في دمج الذكاء الاصطناعي (AI) في الصحة العامة، حيث يتم استخدامه لمجموعة متنوعة من التطبيقات مثل توقع الأمراض، واتخاذ القرار السريري، ومراقبة تفشي الأمراض، والمعلوماتية الصحية. تستند هذه التطبيقات إلى مجموعة واسعة من مصادر البيانات، بما في ذلك السجلات الصحية الإلكترونية ووسائل التواصل الاجتماعي. على الرغم من الفوائد المحتملة، فإن نشر نماذج اللغة الكبيرة (LLMs) في هذا المجال يعيقه مشكلة الهلوسات – المخرجات التي تبدو معقولة ولكنها غير صحيحة من الناحية الواقعية، مما يمكن أن يؤثر سلبًا على التوجيه السريري وصنع السياسات.
تقدم الاستراتيجيات الحالية لمعالجة الهلوسات، مثل التخصيص الدقيق المحدد للمجال، والتدريب العدائي، والتوليد المعزز بالاسترجاع، حلولًا محدودة. ركزت التطورات الأخيرة على اكتشاف الهلوسات من خلال تقنيات مثل الحالات الخفية وخرائط الانتباه، بالإضافة إلى تقطير المعرفة مع تسميات ناعمة لتعزيز الدقة الواقعية أثناء التخصيص الدقيق تحت الإشراف. ومع ذلك، لا يزال هناك فجوة في الأساليب الشاملة المصممة خصيصًا للتخفيف من الهلوسات في سياقات الصحة العامة، مما أدى إلى تطوير MEGA-RAG.
طرق
يتم تقديم إطار MEGA-RAG كنظام جديد للإجابة على الأسئلة المعززة بالاسترجاع (QA) مصمم للقطاعات الطبية الحيوية والصحة العامة، حيث يتناول تحديات الدقة، وقابلية التفسير، وتأسيس الأدلة. على عكس نماذج اللغة الكبيرة التقليدية (LLMs) التي تعتمد فقط على المعرفة الداخلية، يستخدم MEGA-RAG هيكلًا من أربع مراحل لتعزيز الدقة الواقعية وتقليل الهلوسات. يشمل ذلك وحدة استرجاع الأدلة من مصادر متعددة (MSER) لتجميع المعلومات المتنوعة، ووحدة توليد الإجابات المتنوعة (DPAG) لتوليد وإعادة ترتيب إجابات متعددة مرشحة، ووحدة تقييم توافق الأدلة الدلالية (SEAE) لتقييم اتساق الإجابات، ووحدة التوضيح الذاتي المعتمدة على الاختلاف (DISC) لحل الأدلة المتضاربة. يسمح هذا النهج المنظم لـ MEGA-RAG بتقديم استجابات أكثر موثوقية ومتوافقة مع السياسات مقارنةً بالنماذج التقليدية وطرق RAG.
تم استخدام مجموعة بيانات HealthQuestDB، التي تم استخدامها للتدريب والتحقق، وتمت معالجتها بدقة من قبل خبيرين في المجال، محققة اتفاقًا مثاليًا بين المعلقين (κ = 1.0) عبر 20% من الإدخالات. تشمل هذه المجموعة مجموعة واسعة من الموضوعات، بما في ذلك علم الأوبئة وسياسة الصحة، مما يضمن تغطية شاملة وجودة بيانات قوية. تضمنت التقييمات التجريبية نماذج متنوعة، بما في ذلك PubMedBERT وPubMedGPT وLLM + RAG، مع مقاييس الأداء مثل الدقة ودرجة F1. من الجدير بالذكر أن MEGA-RAG أظهر أداءً متفوقًا من خلال تقليل الهلوسات بشكل كبير وتقديم إجابات دقيقة من الناحية الواقعية، كما هو موضح في التحليل المقارن للاستجابات لاستفسار الصحة العامة حول مياه أيونات الفضة. تؤكد هذه النتائج على إمكانية MEGA-RAG في تعزيز صنع السياسات القائمة على الأدلة ودعم اتخاذ القرار السريري في تطبيقات الصحة العامة.
النتائج
في قسم النتائج، يقيم المؤلفون أداء إطار MEGA-RAG في تقليل الهلوسات وتعزيز دقة اتخاذ القرار في السياقات الطبية الحيوية. تم إجراء تجارب مقارنة ضد نماذج أساسية متنوعة، بما في ذلك نماذج محددة للمجال مثل PubMedBERT وPubMedGPT، بالإضافة إلى نماذج اللغة الكبيرة القياسية (LLMs). تشير النتائج إلى أن MEGA-RAG يتفوق على هذه النماذج الأساسية، مما يظهر فعاليته في معالجة التحديات المرتبطة بالهلوسات في التطبيقات الطبية الحيوية.
المناقشة
تسلط قسم المناقشة في ورقة البحث الضوء على التقدم الكبير في دمج رسوم المعرفة ضمن إطار التوليد المعزز بالاسترجاع (RAG)، لا سيما في الطب الحيوي. تشمل المساهمات البارزة KG-RAG، الذي يحسن تضمين الرموز لرسوم المعرفة الطبية الحيوية، وKRAGEN، الذي يعزز قابلية تطبيق RAG في البيئات السريرية. تحسن أساليب أخرى، مثل Zebra-Llama وMedRAG، إمكانية الوصول وموثوقية المعلومات المتعلقة بالأمراض النادرة وتطبيقات الرعاية الصحية، على التوالي. تعالج هذه الابتكارات التحديات مثل الهلوسات في الاستجابات التي ينتجها الذكاء الاصطناعي وتعزز أنظمة اتخاذ القرار السريري.
يتم تقديم إطار MEGA-RAG كوحدة استرجاع أدلة متعددة المصادر تعزز الدقة الواقعية في الإجابة على أسئلة الصحة العامة من خلال الحصول على البيانات من مقالات PubMed، وقاعدة بيانات WHO IRIS، ورسم بياني طبي حيوي منسق. يسهل هذا النهج المنظم التفكير السببي وإزالة الغموض الدلالي، مما يضمن تمثيل العلاقات المعقدة بشكل صريح. يستخدم الإطار كل من الاسترجاع الدلالي الكثيف باستخدام FAISS والاسترجاع المعجمي النادر عبر خوارزمية BM25، مما يسمح باسترجاع فعال بناءً على التشابه المفهومي ومطابقة الكلمات الرئيسية. يتيح دمج ثلاثيات رسوم المعرفة التفكير الصريح حول الروابط الميكانيكية، مما يعزز قوة عملية استرجاع الأدلة. بشكل عام، يظهر MEGA-RAG دقة وموثوقية محسنتين في توليد استجابات قائمة على الأدلة لاستفسارات الصحة العامة، متفوقًا على الهياكل التقليدية لـ RAG.
القيود
يسلط قسم القيود الضوء على عدة اعتبارات عملية تتعلق بمنهجية MEGA-RAG. بينما يظهر النهج أداءً قويًا في تعزيز الدقة، وتوافق الأدلة، والموثوقية لتقديم الإجابات على الأسئلة في مجالات الطب الحيوي والصحة العامة، فإنه يقدم عبئًا حسابيًا إضافيًا بسبب تصميمه متعدد المراحل، الذي يتضمن استرجاعًا متعددًا واستدعاءات لنماذج اللغة الكبيرة. يُعتبر هذا العبء، الذي يؤدي عادةً إلى زيادة زمن الاستجابة لبضع ثوانٍ، مقبولًا في السياقات ذات المخاطر العالية حيث يتم إعطاء الأولوية للدقة الواقعية المحسنة والشفافية.
يشير المؤلفون إلى أن المطالب الحسابية الإضافية تنبع أساسًا من عمليات الاسترجاع متعددة المصادر وإعادة الترتيب، والتي يمكن تنفيذها بالتوازي، مما يتجنب التأخيرات الكبيرة. يعترفون بعدم وجود معايير كفاءة مفصلة في هذه النسخة، حيث أن الكفاءة الحسابية ليست محور التركيز الرئيسي. ومع ذلك، فإنهم يدركون أهمية القابلية للتوسع في البيئات ذات الموارد المحدودة ويقترحون أن البحث المستقبلي يجب أن يستكشف تحسينات على مستوى النظام، بما في ذلك التوازي، واستراتيجيات التخزين المؤقت، واستخدام نماذج المحولات الخفيفة، لتعزيز الاستجابة مع الحفاظ على الموثوقية.
DOI: https://doi.org/10.3389/fpubh.2025.1635381
PMID: https://pubmed.ncbi.nlm.nih.gov/41132171
Publication Date: 2025-10-08
Author(s): Shan Xu et al.
Primary Topic: Topic Modeling
Overview
The research paper presents a novel framework, termed MEGA-RAG (multi-evidence guided answer refinement), aimed at addressing the issue of hallucinations in large language models (LLMs) used in public health. Hallucinations refer to factually inaccurate outputs that can undermine clinical communication and policy-making. MEGA-RAG integrates multi-source evidence retrieval techniques, including dense retrieval via FAISS, keyword-based retrieval via BM, and biomedical knowledge graphs, alongside a cross-encoder reranker for semantic relevance and a discrepancy-aware refinement module to enhance factual accuracy. Experimental results indicate that MEGA-RAG significantly reduces hallucination rates and outperforms baseline models such as PubMedBERT and PubMedGPT across key metrics, including accuracy, precision, recall, and F1 score.
The framework’s modular design not only improves performance but also enhances transparency, allowing users to trace the origins of each assertion made by the model. MEGA-RAG is positioned as a valuable tool for various public health applications, including outbreak surveillance and clinical decision support, by providing evidence-based and contextually relevant responses. Future research directions include expanding evaluations to larger benchmarks, integrating dynamic updates from health agencies, supporting multilingual deployments, and enhancing the self-clarification module with clinician feedback. Additionally, the authors plan to explore efficient deployment strategies suitable for resource-constrained environments, ensuring that the framework remains effective in diverse settings.
Introduction
The introduction highlights the increasing integration of artificial intelligence (AI) in public health, where it is utilized for various applications such as disease prediction, clinical decision-making, outbreak surveillance, and infoveillance. These applications draw on a wide array of data sources, including electronic health records and social media. Despite the potential benefits, the deployment of large language models (LLMs) in this field is hindered by the issue of hallucinations—outputs that appear plausible but are factually incorrect, which can adversely affect clinical guidance and policy-making.
Current strategies to address hallucinations, such as domain-specific fine-tuning, adversarial training, and retrieval-augmented generation, provide limited solutions. Recent advancements have focused on one-shot hallucination detection through techniques like hidden states and attention maps, as well as knowledge distillation with soft labels to enhance factual accuracy during supervised fine-tuning. Nevertheless, there remains a gap in comprehensive methods specifically designed to mitigate hallucinations in public health contexts, which has led to the development of MEGA-RAG.
Methods
The MEGA-RAG framework is introduced as a novel retrieval-augmented question answering (QA) system designed for the biomedical and public health sectors, addressing the challenges of accuracy, interpretability, and evidence grounding. Unlike conventional large language models (LLMs) that rely solely on internal knowledge, MEGA-RAG employs a four-stage architecture to enhance factual accuracy and reduce hallucinations. This includes the Multi-Source Evidence Retrieval (MSER) Module for aggregating diverse information, the Diverse Prompted Answer Generation (DPAG) Module for generating and reranking multiple candidate answers, the Semantic-Evidential Alignment Evaluation (SEAE) Module for assessing answer consistency, and the Discrepancy-Identified Self-Clarification (DISC) Module for resolving conflicting evidence. This structured approach allows MEGA-RAG to provide more reliable and policy-aligned responses compared to traditional LLMs and RAG methods.
The HealthQuestDB dataset, utilized for training and validation, was meticulously annotated by two domain experts, achieving a perfect inter-annotator agreement (Cohen’s κ = 1.0) across 20% of the entries. This dataset encompasses a wide range of topics, including epidemiology and health policy, ensuring comprehensive coverage and robust data quality. Experimental evaluations involved various models, including PubMedBERT, PubMedGPT, and LLM + RAG, with performance metrics such as accuracy and F1-score. Notably, MEGA-RAG demonstrated superior performance by significantly reducing hallucinations and providing factually accurate answers, as illustrated in the comparative analysis of responses to a public health query regarding silver ion water. These results underscore MEGA-RAG’s potential to enhance evidence-based policymaking and clinical decision support in public health applications.
Results
In the Results section, the authors assess the performance of the MEGA-RAG framework in reducing hallucinations and enhancing decision-making accuracy in biomedical contexts. Comparative experiments were conducted against various baseline models, including domain-specific models such as PubMedBERT and PubMedGPT, as well as standard large language models (LLMs). The findings indicate that MEGA-RAG outperforms these baselines, demonstrating its effectiveness in addressing the challenges associated with hallucinations in biomedical applications.
Discussion
The discussion section of the research paper highlights significant advancements in integrating knowledge graphs within the Retrieval-Augmented Generation (RAG) framework, particularly in biomedicine. Notable contributions include KG-RAG, which optimizes token embedding for biomedical knowledge graphs, and KRAGEN, which enhances RAG’s applicability in clinical settings. Other approaches, like Zebra-Llama and MedRAG, further improve the accessibility and reliability of information related to rare diseases and healthcare applications, respectively. These innovations address challenges such as hallucinations in AI-generated responses and bolster clinical decision-making systems.
The MEGA-RAG framework is introduced as a multi-source evidence retrieval module that enhances factual accuracy in public health question answering by concurrently sourcing data from PubMed articles, the WHO IRIS database, and a curated biomedical knowledge graph. This structured approach facilitates causal reasoning and semantic disambiguation, ensuring that complex relationships are explicitly represented. The framework employs both dense semantic retrieval using FAISS and sparse lexical retrieval via the BM25 algorithm, allowing for effective retrieval based on conceptual similarity and keyword matching. The integration of knowledge graph triples enables explicit reasoning over mechanistic associations, enhancing the robustness of the evidence retrieval process. Overall, MEGA-RAG demonstrates improved accuracy and reliability in generating evidence-based responses for public health queries, outperforming traditional RAG architectures.
Limitations
The section on limitations highlights several practical considerations regarding the MEGA-RAG methodology. While the approach demonstrates strong performance in enhancing accuracy, evidence alignment, and robustness for biomedical and public health question-answering (QA), it introduces additional computational overhead due to its multi-stage design, which involves multiple retrieval and large language model calls. This overhead, typically resulting in latency increases of a few seconds, is deemed acceptable in high-stakes contexts where improved factual fidelity and transparency are prioritized.
The authors note that the additional computational demands primarily stem from multi-source retrieval and reranking processes, which can be parallelized, thus avoiding significant delays. They acknowledge the absence of detailed efficiency benchmarks in this version, as computational efficiency is not the main focus. However, they recognize the importance of scalability in resource-constrained environments and suggest that future research should investigate system-level optimizations, including parallelization, caching strategies, and the use of lightweight transformer models, to enhance responsiveness while maintaining reliability.
