DOI: https://doi.org/10.1186/s12911-024-02459-6
PMID: https://pubmed.ncbi.nlm.nih.gov/38475802
تاريخ النشر: 2024-03-12
المؤلف: Ye‐Jean Park وآخرون
الموضوع الرئيسي: الذكاء الاصطناعي في الرعاية الصحية والتعليم
نظرة عامة
تستعرض هذه المراجعة نطاق دور نماذج اللغة الكبيرة (LLMs)، مثل ChatGPT من OpenAI، في التطبيقات السريرية، مع تسليط الضوء على فوائدها المحتملة وقيودها الجوهرية. تلخص المراجعة النتائج من 55 دراسة، كاشفة أن LLMs يمكن أن تساعد بفعالية في مهام مثل تجميع ملاحظات المرضى والمساعدة في التنقل في الرعاية الصحية، بينما تدعم أيضًا اتخاذ القرارات السريرية عند استخدامها جنبًا إلى جنب مع الإشراف البشري. ومع ذلك، لا تزال التحديات قائمة، بما في ذلك التحيزات في بيانات التدريب، وخطر توليد معلومات مضللة، ومخاوف أخلاقية وقانونية واجتماعية واقتصادية متنوعة قد تؤثر على رعاية المرضى.
تؤكد المراجعة على ضرورة وجود إطار تقييم موحد لتقييم الفائدة السريرية لـ LLMs، حيث تفتقر الطرق الحالية إلى الاتساق. تدعو إلى دراسات التحقق والتكرار لضمان دقة وموثوقية LLMs قبل دمجها على نطاق واسع في بيئات الرعاية الصحية. يدعو المؤلفون إلى نهج تعاوني يشمل العديد من أصحاب المصلحة، وخاصة الباحثين، لمعالجة القضايا اللوجستية والأخلاقية المستمرة. في النهاية، بينما تقدم LLMs فرصًا كبيرة لتعزيز تقديم الرعاية الصحية، فإن الاعتبار الدقيق لقيودها أمر ضروري لتعظيم فوائدها وتقليل المخاطر المحتملة.
مقدمة
تناقش مقدمة هذه الورقة البحثية قدرات وتطبيقات نماذج اللغة الكبيرة (LLMs)، مع التركيز بشكل خاص على نموذج OpenAI’s Generative Pre-training Transformer (GPT) وواجهته، ChatGPT. تستخدم هذه النماذج مجموعات نصوص غير مصنفة واسعة النطاق وتعلم التعزيز من ردود الفعل البشرية لتعلم الأنماط اللغوية والعلاقات السياقية، مما يمكنها من توليد استجابات متماسكة تشبه الإنسان لمجموعة متنوعة من المدخلات. يمثل ChatGPT، الذي تم إطلاقه في نوفمبر 2022، مثالاً على هذه التكنولوجيا، حيث تم تدريبه على مصادر متنوعة بما في ذلك CommonCrawl وWikipedia. تتيح قدرة النموذج على التعلم المستمر له تحسين استجاباته بمرور الوقت من خلال الاحتفاظ بذاكرة التفاعلات السابقة.
تسلط الورقة الضوء على التطبيقات المحتملة لـ LLMs في الرعاية الصحية، مثل تعزيز دعم اتخاذ القرارات السريرية وتبسيط المهام الإدارية مثل تسجيل بيانات المرضى وتوليد ملاحظات الخروج. يمكن أن تحسن هذه القدرات الكفاءة وتسمح لمقدمي الرعاية الصحية بالتركيز أكثر على رعاية المرضى. ومع ذلك، يشير المؤلفون أيضًا إلى تحديات كبيرة، بما في ذلك التحيزات في بيانات التدريب، وخطر توليد معلومات غير دقيقة (يشار إليها باسم “الهلاوس”)، والمخاوف الأخلاقية المتعلقة بسلامة المرضى وخصوصيتهم. نظرًا لعدم وجود تلخيص كافٍ للمعرفة الحالية حول التطبيقات السريرية لـ LLMs حتى يونيو 2023، أجرى المؤلفون مراجعة نطاقية لتلخيص وتنظيم الرؤى من الدراسات الحديثة، بهدف إبلاغ المناقشات المستقبلية حول دمج LLMs في البيئات السريرية.
طرق البحث
في هذا القسم، يوضح المؤلفون منهجيتهم لإجراء مراجعة نطاقية تهدف إلى تقييم الفائدة السريرية لنماذج اللغة الكبيرة (LLMs) في بيئات الرعاية الصحية. وفقًا لإرشادات PRISMA-ScR وإطار عمل Arksey وO’Malley، شملت المراجعة 36 مقالة رئيسية، تضم دراسات تجريبية ومطبوعات مسبقة. قام المؤلفون بتطوير بروتوكول مراجعة تم تسجيله علنًا على إطار العلوم المفتوحة. قاموا باستخراج وتحليل السمات الرئيسية من هذه الدراسات، مع التركيز على كيفية تحفيز الباحثين لـ LLMs مثل ChatGPT وتقييم الاستجابات.
تكشف النتائج عن مجموعة من تقنيات التحفيز المستخدمة في الدراسات، من التحفيزات الرسمية والمنظمة إلى الاستفسارات غير الرسمية والمحادثات. على سبيل المثال، طلب الباحثون من ChatGPT توليد رسائل للمرضى أو نماذج موافقة جراحية بناءً على سيناريوهات سريرية محددة. غالبًا ما شملت تقييم الاستجابات مراجعين بشريين يقيمون الدقة وقابلية القراءة، وأحيانًا ما تم دعمها بأدوات عبر الإنترنت. كما سلطت المراجعة الضوء على التباين في صرامة طرق التحفيز والتقييم، مما يثير تساؤلات حول ما يشكل استجابة “دقيقة” أو “قابلة للقراءة”. يعقد هذا التباين تفسير النتائج، خاصة في سياق التطبيقات الطبية، حيث تظل المعايير لتقييم مخرجات LLMs غامضة.
النتائج
أسفرت استراتيجية البحث عن إجمالي 4,036 مقالة، تم من خلالها إزالة 998 مكررًا من خلال عمليات الفحص التلقائي واليدوي. بعد فحص 3,126 عنوانًا وملخصًا، تم اعتبار 63 مقالة مؤهلة للمراجعة الكاملة للنص. ومع ذلك، تم استبعاد 8 من هذه المقالات بسبب عدم توافقها مع معايير الإدراج المحددة، مما أسفر عن 55 مقالة تم تضمينها في التحليل في النهاية.
لضمان جودة وصرامة استخراج البيانات، نفذ المؤلفون عملية استخراج موحدة استنادًا إلى مثال متفق عليه، إلى جانب فحوصات وتأكيدات مكررة. تم رسم البيانات ذات الصلة من كل مقالة مدرجة بالنسبة لأسئلة وأهداف المراجعة، مما يسهل تلخيص منهجي للنتائج التي تعالج أهداف الدراسة.
المناقشة
تتناول قسم المناقشة في الورقة البحثية المنهجية والنتائج المتعلقة بالجدوى السريرية لنماذج اللغة الكبيرة (LLMs) في بيئات الرعاية الصحية. تم إجراء بحث شامل في قواعد البيانات عبر MEDLINE وEMBASE وCINAHL، مع التركيز على المنشورات باللغة الإنجليزية التي تستكشف استخدام LLMs في البيئات السريرية. شملت الدراسة أنواعًا مختلفة من المنشورات، مثل المقالات البحثية والمطبوعات المسبقة والتعليقات، مع استبعاد تلك التي لا تتعلق مباشرة بالتطبيقات السريرية. صنفت عملية استخراج البيانات الدراسات إلى مواضيع، بما في ذلك فائدة LLMs في توثيق المرضى، والاعتبارات الأخلاقية واللوجستية، ودعم تنقل المرضى، وعمليات اتخاذ القرار السريري.
تشير النتائج الرئيسية إلى أن LLMs، وخاصة ChatGPT، يمكن أن تعزز الكفاءة في تجميع ملاحظات المرضى وتحسين التواصل مع المرضى من خلال تبسيط المعلومات الطبية المعقدة. ومع ذلك، تم تسليط الضوء أيضًا على التحديات مثل عدم الدقة في التوثيق، والتحيزات المحتملة، والمخاوف الأخلاقية المتعلقة بخصوصية البيانات. تؤكد الورقة على الحاجة إلى مزيد من البحث لتأسيس أطر تقييم موحدة لـ LLMs، ومعالجة الآثار الأخلاقية والقانونية، وإشراك مجموعة متنوعة من أصحاب المصلحة في عملية الدمج. تشمل التوصيات تطوير تطبيقات LLM محلية لتعزيز الخصوصية والامتثال للوائح، بالإضافة إلى تعزيز التعاون بين محترفي الرعاية الصحية وصانعي السياسات ومطوري التكنولوجيا لضمان التنفيذ الفعال والإشراف على LLMs في الممارسة السريرية.
القيود
استخدمت المراجعة النطاقية المقدمة في هذه الورقة منهجيات صارمة، ملتزمة بقائمة مراجعة PRISMA-ScR واستخدام بروتوكول منشور مسبقًا تم تطويره بواسطة خبراء في تجميع الأدلة. تم إجراء عمليات بحث شاملة عبر ثلاث قواعد بيانات رئيسية، مدعومة بعمليات بحث يدوية موسعة، مما يضمن شمولية الأدبيات ذات الصلة. ومع ذلك، على الرغم من هذه الجهود، تعترف المراجعة بالقيود المحتملة في استراتيجيتها البحثية، التي استخدمت ثمانية مصطلحات رئيسية تتعلق بنماذج اللغة الكبيرة (LLMs). قد يكون هذا النهج قد استبعد عن غير قصد دراسات ذات صلة لم تتضمن مصطلحات مرتبطة بالصحة في عناوينها أو كلماتها الرئيسية.
علاوة على ذلك، فإن المشهد المتطور بسرعة لـ LLMs يشكل تحديًا، حيث أن التحليل يعتمد بشكل أساسي على الدراسات المتاحة في وقت البحث، مما قد يتسبب في إغفال التطورات الأكثر حداثة. كما تشير المراجعة إلى غياب أداة تقييم جودة موثوقة لدراسات التعلم الآلي الاستكشافية، مما يحد من القدرة على تقييم خطر التحيز بين المقالات المدرجة. بالإضافة إلى ذلك، لا يمكن تأكيد دقة وموثوقية النتائج من المقالات المسبقة المدرجة في المراجعة، حيث لم تخضع هذه المقالات لمراجعة الأقران.
DOI: https://doi.org/10.1186/s12911-024-02459-6
PMID: https://pubmed.ncbi.nlm.nih.gov/38475802
Publication Date: 2024-03-12
Author(s): Ye‐Jean Park et al.
Primary Topic: Artificial Intelligence in Healthcare and Education
Overview
This scoping review investigates the role of large language models (LLMs), such as OpenAI’s ChatGPT, in clinical applications, highlighting their potential benefits and inherent limitations. The review synthesizes findings from 55 studies, revealing that LLMs can effectively assist in tasks like compiling patient notes and aiding healthcare navigation, while also supporting clinical decision-making when used alongside human oversight. However, challenges persist, including biases in training data, the risk of generating misleading information, and various ethical, legal, and socioeconomic concerns that may impact patient care.
The review emphasizes the necessity for a standardized evaluation framework to assess the clinical utility of LLMs, as current methods lack consistency. It calls for validation and replication studies to ensure the accuracy and reliability of LLMs before their widespread integration into healthcare settings. The authors advocate for a collaborative approach involving multiple stakeholders, particularly researchers, to address ongoing logistical and ethical issues. Ultimately, while LLMs present significant opportunities for enhancing healthcare delivery, careful consideration of their limitations is essential for maximizing their benefits and minimizing potential risks.
Introduction
The introduction of this research paper discusses the capabilities and applications of large language models (LLMs), particularly focusing on OpenAI’s Generative Pre-training Transformer (GPT) and its interface, ChatGPT. These models utilize extensive unlabelled text corpora and reinforcement learning from human feedback to learn linguistic patterns and contextual relationships, enabling them to generate coherent, human-like responses to various inputs. ChatGPT, launched in November 2022, exemplifies this technology, having been trained on diverse sources including CommonCrawl and Wikipedia. The model’s continual learning ability allows it to refine its responses over time by retaining memory of previous interactions.
The paper highlights the potential applications of LLMs in healthcare, such as enhancing clinical decision support and streamlining administrative tasks like patient charting and discharge note generation. These capabilities could improve efficiency and allow healthcare providers to concentrate more on patient care. However, the authors also note significant challenges, including biases in training data, the risk of generating inaccurate information (referred to as “hallucination”), and ethical concerns regarding patient safety and privacy. Given the limited synthesis of existing knowledge on the clinical applications of LLMs as of June 2023, the authors conducted a scoping review to summarize and organize insights from recent studies, aiming to inform future discussions on the integration of LLMs in clinical settings.
Methods
In this section, the authors outline their methodology for conducting a scoping review aimed at assessing the clinical utility of large language models (LLMs) in healthcare settings. Following the PRISMA-ScR guidelines and the Arksey and O’Malley framework, the review included 36 primary articles, encompassing empirical studies and preprints. The authors developed a review protocol that was publicly registered on the Open Science Framework. They extracted and analyzed key traits from these studies, focusing on how researchers prompted LLMs like ChatGPT and evaluated the responses.
The findings reveal a range of prompting techniques used in the studies, from formal, structured prompts to informal, conversational queries. For instance, researchers asked ChatGPT to generate patient letters or surgical consent forms based on specific clinical scenarios. Evaluation of the responses often involved human reviewers assessing accuracy and readability, sometimes supplemented by online tools. The review also highlighted variability in the rigor of prompting and evaluation methods, raising questions about what constitutes an “accurate” or “readable” response. This variability complicates the interpretation of results, particularly in the context of medical applications, where the standards for evaluating LLM outputs remain ambiguous.
Results
The search strategy yielded a total of 4,036 articles, from which 998 duplicates were removed through both automatic and manual screening processes. Following the screening of 3,126 titles and abstracts, 63 articles were deemed eligible for full-text review. However, 8 of these were excluded due to misalignment with the established inclusion criteria, resulting in 55 articles that were ultimately included in the analysis.
To ensure the quality and rigor of data extraction, the authors implemented a standardized extraction process based on a consensus example, alongside duplicate screenings and confirmations. The relevant data from each included article were charted in relation to the review questions and objectives, facilitating a systematic synthesis of results that addresses the study’s aims.
Discussion
The discussion section of the research paper outlines the methodology and findings related to the clinical applicability of large language models (LLMs) in healthcare settings. A comprehensive database search was conducted across MEDLINE, EMBASE, and CINAHL, focusing on English-language publications that explore LLMs’ use in clinical environments. The study included various types of publications, such as research articles, preprints, and commentaries, while excluding those that did not pertain directly to clinical applications. The data extraction categorized studies into themes, including LLMs’ utility in patient documentation, ethical and logistical considerations, patient navigation support, and clinical decision-making processes.
Key findings indicate that LLMs, particularly ChatGPT, can enhance efficiency in compiling patient notes and improve communication with patients by simplifying complex medical information. However, challenges such as inaccuracies in documentation, potential biases, and ethical concerns regarding data privacy were also highlighted. The paper emphasizes the need for further research to establish standardized evaluation frameworks for LLMs, address ethical and legal implications, and involve diverse stakeholders in the integration process. Recommendations include developing localized LLM applications to enhance privacy and compliance with regulations, as well as fostering collaboration among healthcare professionals, policymakers, and technology developers to ensure effective implementation and oversight of LLMs in clinical practice.
Limitations
The scoping review presented in this paper employed rigorous methodologies, adhering to the PRISMA-ScR checklist and utilizing a prospectively published protocol developed by experts in evidence synthesis. Comprehensive searches were conducted across three major databases, supplemented by extensive handsearching, ensuring a thorough inclusion of relevant literature. However, despite these efforts, the review acknowledges potential limitations in its search strategy, which utilized eight key terms related to large language models (LLMs). This approach may have inadvertently excluded pertinent studies that did not incorporate health-related terminology in their titles or keywords.
Moreover, the rapidly evolving landscape of LLMs poses a challenge, as the analysis is primarily based on studies available at the time of the search, potentially omitting more recent developments. The review also notes the absence of a validated quality assessment tool for exploratory machine learning studies, which limits the ability to evaluate the risk of bias among the included articles. Additionally, the accuracy and validity of findings from preprint articles included in the review cannot be confirmed, as these have not undergone peer review.
