تنفيذ نماذج اللغة الكبيرة في السجلات الصحية الإلكترونية
Implementation of large language models in electronic health records

شارك:
المجلة: PLOS Digital Health، المجلد: 4، العدد: 12
DOI: https://doi.org/10.1371/journal.pdig.0001141
PMID: https://pubmed.ncbi.nlm.nih.gov/41417848
تاريخ النشر: 2025-12-19
المؤلف: Maxime Griot وآخرون
الموضوع الرئيسي: أنظمة السجلات الصحية الإلكترونية

نظرة عامة

كان الهدف من تنفيذ مساعد نموذج اللغة الكبير (LLM) المتوافق مع GDPR والمُدمج في نظام السجلات الصحية الإلكترونية (EHR) في مستشفى جامعي أوروبي هو تخفيف عبء التوثيق على الأطباء. باستخدام نموذج Qwen3-235B مع توليد معزز بالاسترجاع، قدم المساعد استجابات مدركة للسياق بناءً على بيانات المرضى المنظمة، والمستندات السريرية، والأدبيات الطبية. أظهر اختبار تجريبي لمدة شهر شمل 28 طبيبًا عبر تسع تخصصات تفاعلًا كبيرًا، حيث استخدم 64% من المشاركين المساعد يوميًا وولدوا 482 محادثة متعددة الأدوار. كانت المهام الرئيسية تشمل التلخيص، واسترجاع المعلومات، وصياغة الملاحظات، والتي شكلت مجتمعة أكثر من 70% من التفاعلات.

بعد التجربة، تم نشر مساعد LLM في جميع أنحاء المستشفى، حيث وصل إلى 1,028 مستخدمًا شاركوا في 14,910 محادثات على مدى خمسة أشهر، مع استخدام أكثر من نصف الأطباء له على الأقل أسبوعيًا. أشارت أنماط الاستخدام إلى تكامل مستقر في سير العمل السريري، يركز بشكل أساسي على الوصول إلى المعلومات ودعم التوثيق. ومن الجدير بالذكر أن حجم التعليقات انخفض بعد التجربة، مما يشير إلى أن الاستخدام الروتيني قد يقلل من الإبلاغ الطوعي، مما يبرز ضرورة استراتيجيات المراقبة الآلية. بشكل عام، تُظهر الدراسة أن التكامل على نطاق واسع لنماذج LLM في البيئات السريرية ممكن تقنيًا ويمكن أن يحافظ على الاستخدام المستدام عند دمجه بشكل فعال ضمن سير العمل في السجلات الصحية الإلكترونية، مما يعزز الإمكانية لنشر الذكاء الاصطناعي المسؤول في الرعاية الصحية.

مقدمة

لقد حول إدخال السجلات الصحية الإلكترونية (EHRs) بشكل كبير المجال الطبي من خلال تعزيز استمرارية الرعاية وتوحيد الوصول إلى البيانات، بينما تولد أيضًا مجموعات بيانات أكبر للبحث. ومع ذلك، تأتي هذه الرقمنة مع عيوب، لا سيما زيادة عبء التوثيق على الأطباء، الذين قد يقضون أكثر من 50% من وقتهم في مهام متعلقة بالسجلات الصحية الإلكترونية. ويرجع ذلك إلى ضرورة إدخال نفس المعلومات في تنسيقات متعددة، مما يؤدي إلى إما وقت توثيق مفرط أو إدخالات غير مكتملة في الحقول المنظمة. ونتيجة لذلك، يواجه الأطباء ضغطًا متزايدًا ووقت رعاية المرضى منخفض، حيث يقضي أطباء الطوارئ في المتوسط 20 دقيقة من التحضير لكل مريض بسبب الحجم الهائل من المستندات السريرية.

استجابةً لهذه التحديات، ظهرت نماذج اللغة الكبيرة (LLMs) كحل واعد، حيث أظهرت أداءً قويًا على المعايير الطبية وقدرة على معالجة النصوص الكبيرة بكفاءة. ومع ذلك، تركزت الغالبية العظمى من الأبحاث الحالية على مهام التفكير السريري مثل التشخيص والتخطيط للعلاج ضمن بيئات اصطناعية، بدلاً من التطبيقات الواقعية. تعقد صعوبة الوصول إلى بيانات سريرية حقيقية ترجمة قدرات LLM إلى استخدام عملي، حيث تعتمد أدوات التقييم الحالية بشكل أساسي على أسئلة الاختيار المتعدد (MCQs) التي قد لا تلتقط تمامًا تعقيدات الممارسة السريرية. بينما يتم الاعتراف بأهمية الاختبار في العالم الحقيقي في الأدبيات الطبية، لا يزال هناك فجوة ملحوظة في التقييمات الشاملة، والتي غالبًا ما تقتصر على مهام محددة في بيئات خاضعة للرقابة. تشمل الطريقتان السائدتان للتقييم السريري في العالم الحقيقي إما توليد أو تحويل بيانات نصية موجودة أو تحويل بيانات الكلام إلى تنسيقات نصية.

طرق

في هذا القسم، يوضح المؤلفون الطرق المستخدمة لتطوير نظام يهدف إلى الاندماج بسلاسة في سير العمل اليومي للأطباء. تم إجراء ورش عمل أولية لتثقيف الأطباء حول نماذج اللغة الكبيرة (LLMs) وجمع آرائهم حول حالات الاستخدام المحتملة. أسفر هذا النهج التعاوني عن قائمة ذات أولوية لحالات الاستخدام، مع تركيز قوي على تعزيز الوصول إلى أنواع مختلفة من المعلومات، بما في ذلك بيانات المرضى، والأدبيات العلمية، والبروتوكولات الداخلية. الحل المقترح هو دردشة آلية قادرة على استرجاع وتلخيص والتفاعل مع هذه المصادر المعلوماتية، مما يعالج الغالبية العظمى من الاحتياجات المحددة.

لضمان اعتماد الأداة بشكل فعال دون إضافة عبء عمل على الأطباء، تم تحديد ثلاث استراتيجيات رئيسية: (1) يجب استضافة التطبيق داخل نظام السجلات الصحية الإلكترونية (EHR) لتقليل تبديل السياقات وتوحيد المعلومات؛ (2) يجب أن تكون واجهة المستخدم بديهية ومألوفة، مستفيدة من مبادئ التصميم للأدوات المستخدمة على نطاق واسع مثل ChatGPT لتسهيل الاستخدام دون الحاجة إلى تدريب مكثف؛ و(3) يجب أن يكون التكامل غير مزعج، يظهر فقط عند الحاجة لتجنب تعطيل سير العمل الحالي.

النتائج

شملت دراسة الطيار التي استمرت شهرًا 28 طبيبًا استشاريًا كبيرًا عبر تسع تخصصات، بما في ذلك الأورام، وكبار السن، والطب الباطني، وطب الأطفال، والعناية المركزة، والجراحة، وطب الطوارئ، وطب العيون، والعلاج الإشعاعي. شكلت التخصصات المشاركة حجمًا كبيرًا من حالات الاستشفاء والتدخلات في عام 2024، مع إجمالي 35,354 حالة استشفاء و20,626 تدخل جراحي، بالإضافة إلى 79,506 زيارات طوارئ. كان التركيز الأساسي للتدخل على أنماط ممارسة الأطباء، دون تقييم نتائج سريرية على مستوى المرضى.

في البداية، سمحت الدراسة باسترجاع المستندات السريرية الداخلية لتعريف الأطباء بالنظام، مما قلل من التعقيد خلال المرحلة المبكرة. في الأسبوع الأخير من التجربة، تم تفعيل النظام الكامل، مما أتاح تقييمًا شاملاً لتأثير التدخل على الممارسة السريرية.

مناقشة

تسلط المناقشة الضوء على الدور المتطور لنماذج اللغة الكبيرة (LLMs) في سير العمل السريري، مع التأكيد على قدرتها على تعزيز التوثيق والتفكير التشخيصي. تشير الدراسات، مثل تلك التي أجراها ديكر وآخرون وزاريتسكي وآخرون، إلى أن المستندات التي تم إنشاؤها بواسطة LLM، بما في ذلك نماذج الموافقة المستنيرة وملخصات الخروج، يمكن أن تتطابق أو تتجاوز جودة تلك التي ينتجها الأطباء البشر. ومع ذلك، لا تزال هناك تحديات، حيث وجدت تاي-سيل وآخرون عدم وجود تقليل كبير في عبء العمل من الردود الآلية على رسائل المرضى، وأشار باكستر وآخرون إلى أن مسودات LLM غالبًا ما تتطلب مراجعات كبيرة، خاصةً للتواصل الحساس. أظهر فريدولجاك وآخرون أيضًا أن بعض نماذج LLM يمكن أن تحقق دقة تشخيصية عالية في السيناريوهات الواقعية، مما يبرز قدرتها على دعم اتخاذ القرار السريري، على الرغم من أن أداء النموذج يختلف بشكل كبير بناءً على الحجم والتدريب.

أظهر دمج تقنيات تحويل الكلام إلى نص، وخاصة الكتبة الذكية المحيطة، وعدًا في تقليل أعباء التوثيق وتحسين رضا الأطباء. تشير الدراسات التجريبية، مثل تلك التي أجراها بالوك وآخرون وكايزر بيرمانينتي، إلى توفير كبير في الوقت وزيادة في تفاعل المرضى، على الرغم من أن الحاجة لمراجعة الأطباء للملاحظات التي تم إنشاؤها بواسطة الذكاء الاصطناعي لا تزال قائمة. تؤكد الورقة على أنه بينما تظهر نماذج LLM كفاءة قابلة للقياس في المهام السريرية، فإن التحقق الصارم وإطارات الحوكمة ضرورية للنشر الآمن. يقترح المؤلفون نموذجًا شاملاً للأمان والحوكمة لدمج LLM في أنظمة السجلات الصحية الإلكترونية (EHR)، مع التركيز على اعتماد المستخدم وأنماط التفاعل بدلاً من النتائج السريرية. يهدف هذا النهج إلى معالجة التحديات التقنية والأخلاقية لتنفيذ أدوات الذكاء الاصطناعي في الرعاية الصحية، مع ضمان الامتثال للوائح الخصوصية وتعزيز ثقة الأطباء في سير العمل المدعوم بالذكاء الاصطناعي.

Journal: PLOS Digital Health, Volume: 4, Issue: 12
DOI: https://doi.org/10.1371/journal.pdig.0001141
PMID: https://pubmed.ncbi.nlm.nih.gov/41417848
Publication Date: 2025-12-19
Author(s): Maxime Griot et al.
Primary Topic: Electronic Health Records Systems

Overview

The implementation of a secure, GDPR-compliant Large Language Model (LLM) assistant integrated into the Epic Electronic Health Record (EHR) system at a European university hospital aimed to alleviate the documentation burden on clinicians. Utilizing the Qwen3-235B model with Retrieval Augmented Generation, the assistant provided context-aware responses based on structured patient data, clinical documents, and medical literature. A one-month pilot involving 28 physicians across nine specialties revealed significant engagement, with 64% of participants using the assistant daily and generating 482 multi-turn conversations. The primary tasks included summarization, information retrieval, and note drafting, which collectively accounted for over 70% of interactions.

Following the pilot, the LLM assistant was deployed hospital-wide, reaching 1,028 users who engaged in 14,910 conversations over five months, with more than half of the clinicians using it at least weekly. The usage patterns indicated a stable integration into clinical workflows, primarily focused on information access and documentation support. Notably, the volume of feedback decreased after the pilot, suggesting that routine use may reduce voluntary reporting, highlighting the necessity for automated monitoring strategies. Overall, the study demonstrates that large-scale integration of LLMs into clinical settings is technically viable and can maintain sustained usage when effectively embedded within EHR workflows, reinforcing the potential for responsible AI deployment in healthcare.

Introduction

The introduction of Electronic Health Records (EHRs) has significantly transformed the medical field by enhancing continuity of care and centralizing data access, while also generating larger datasets for research. However, this digitization comes with drawbacks, notably the increased documentation burden on physicians, who may spend over 50% of their time on EHR-related tasks. This is largely due to the necessity of entering the same information in multiple formats, leading to either excessive documentation time or incomplete entries in structured fields. Consequently, clinicians face heightened stress and reduced patient care time, with emergency physicians averaging 20 minutes of preparation per patient due to the overwhelming volume of clinical documents.

In response to these challenges, large language models (LLMs) have emerged as a promising solution, demonstrating strong performance on medical benchmarks and the ability to process extensive text efficiently. However, the majority of existing research has concentrated on clinical reasoning tasks such as diagnostics and treatment planning within synthetic environments, rather than real-world applications. The difficulty of accessing authentic clinical data complicates the translation of LLM capabilities into practical use, with current assessment tools primarily relying on multiple-choice questions (MCQs) that may not fully capture the complexities of clinical practice. While the significance of real-world testing is acknowledged in medical literature, there remains a notable gap in comprehensive evaluations, which are often limited to specific tasks in controlled settings. The two prevalent methods for real-world clinical evaluation involve either generating or transforming existing textual data or converting speech data into textual formats.

Methods

In this section, the authors outline the methods employed to develop a system aimed at integrating seamlessly into physicians’ daily workflows. Initial workshops were conducted to educate physicians about large language models (LLMs) and gather their input on potential use cases. This collaborative approach resulted in a prioritized list of use cases, with a strong emphasis on enhancing access to various types of information, including patient data, scientific literature, and internal protocols. A proposed solution is a chatbot capable of retrieving, summarizing, and interacting with these information sources, thereby addressing the majority of identified needs.

To ensure the tool’s effective adoption without adding to physicians’ workloads, three key strategies were identified: (1) the application should be hosted within the Electronic Health Record (EHR) system to minimize context switching and centralize information; (2) the user interface must be intuitive and familiar, leveraging the design principles of widely used tools like ChatGPT to facilitate ease of use without requiring extensive training; and (3) the integration should be unobtrusive, appearing only when needed to avoid disrupting existing workflows.

Results

The one-month pilot study involved 28 senior attending physicians across nine specialties, including oncology, geriatrics, internal medicine, pediatrics, intensive care, surgery, emergency medicine, ophthalmology, and radiotherapy. The participating specialties accounted for a significant volume of hospitalizations and interventions in 2024, with a total of 35,354 hospitalizations and 20,626 surgical interventions, alongside 79,506 emergency visits. The primary focus of the intervention was on the practice patterns of physicians, with no patient-level clinical outcomes being assessed.

Initially, the study allowed for internal clinical document retrieval to familiarize clinicians with the system, minimizing complexity during the early phase. In the final week of the pilot, the complete system was activated, enabling a comprehensive evaluation of the intervention’s impact on clinical practice.

Discussion

The discussion highlights the evolving role of large language models (LLMs) in clinical workflows, emphasizing their potential to enhance documentation and diagnostic reasoning. Studies, such as those by Decker et al. and Zaretsky et al., indicate that LLM-generated documents, including informed consent forms and discharge summaries, can match or exceed the quality of those produced by human clinicians. However, challenges remain, as Tai-Seale et al. found no significant workload reduction from automated patient-message replies, and Baxter et al. noted that LLM drafts often require substantial revisions, particularly for sensitive communications. Vrdoljak et al. further demonstrated that certain LLMs can achieve high diagnostic accuracy in real-world scenarios, underscoring their potential to support clinical decision-making, although model performance varies significantly based on scale and training.

The integration of speech-to-text technologies, particularly ambient AI scribes, has shown promise in reducing documentation burdens and improving clinician satisfaction. Pilot studies, such as those conducted by Balloch et al. and Kaiser Permanente, report significant time savings and enhanced patient engagement, although the need for clinician review of AI-generated notes persists. The paper emphasizes that while LLMs demonstrate measurable competence in clinical tasks, rigorous validation and governance frameworks are essential for safe deployment. The authors propose a comprehensive security and governance model for LLM integration into electronic health record (EHR) systems, focusing on user adoption and interaction patterns rather than clinical outcomes. This approach aims to address the technical and ethical challenges of implementing AI tools in healthcare, ensuring compliance with privacy regulations and fostering clinician trust in AI-assisted workflows.

شارك: