استخراج معلومات البروتوكول المدعوم بالذكاء الاصطناعي لتحسين الدقة والكفاءة في سير عمل التجارب السريرية
AI-assisted protocol information extraction for improved accuracy and efficiency in clinical trial workflows

شارك:
المجلة: Journal of Biomedical Informatics، المجلد: 179
DOI: https://doi.org/10.1016/j.jbi.2026.105036
PMID: https://pubmed.ncbi.nlm.nih.gov/41967790
تاريخ النشر: 2026-04-10
المؤلف: Ramtin Babaeipour وآخرون
الموضوع الرئيسي: الحوسبة العلمية وإدارة البيانات

نظرة عامة

تسبب التعقيد المتزايد لبروتوكولات التجارب السريرية، إلى جانب التعديلات المتكررة والتحديات في إدارة المعرفة، أعباء كبيرة على فرق التجارب. تقيم هذه الدراسة نظام الذكاء الاصطناعي (AI) الذي يستخدم نماذج اللغة الكبيرة التوليدية (LLMs) مع توليد معزز بالاسترجاع (RAG) للاستخراج الآلي لمعلومات بروتوكول التجارب السريرية. تقارن الأبحاث دقة الاستخراج لعملية RAG الخاصة بالتجارب السريرية، التي حققت دقة بنسبة 89.0%، مقابل LLMs المستقلة التي حققت دقة بنسبة 62.6% عند ضبطها مع المطالبات، باستخدام ملاحظات مرجعية مدعومة من الخبراء للتحقق.

بالإضافة إلى ذلك، تم تقييم التأثير التشغيلي لمساعدة الذكاء الاصطناعي على سير العمل المحاكي لمنسقي الأبحاث السريرية (CRCs). أظهرت النتائج أن المهام المدعومة بالذكاء الاصطناعي تم إنجازها بسرعة لا تقل عن 40%، وتم اعتبارها أقل تطلبًا من الناحية المعرفية، وحصلت على تفضيل قوي من المستخدمين. بينما يظل الإشراف من الخبراء أمرًا حيويًا، تشير هذه النتائج إلى أن الاستخراج المدعوم بالذكاء الاصطناعي يمكن أن يسهل ذكاء البروتوكول على نطاق واسع. وهذا يحفز دمج مثل هذه المنهجيات في سير العمل السريري في العالم الحقيقي لاستكشاف آثارها على الجدوى، وبدء الدراسة، ومراقبة ما بعد التفعيل.

مقدمة

تسلط مقدمة هذه الورقة البحثية الضوء على الدور الحاسم لبروتوكولات التجارب السريرية في ضمان التدخلات الطبية الفعالة والآمنة. تؤكد أن هذه البروتوكولات تعمل كأدوات تواصل أساسية بين المحققين، وفرق البحث، والمشاركين، والمجتمع العلمي، موضحة خلفية التجربة، وأهدافها، وتصميمها. ومع ذلك، أدى التعقيد المتزايد للبروتوكولات إلى تباين في اكتمالها والامتثال لإرشادات الجودة، مما أدى إلى تعديلات متكررة تستغرق وقتًا طويلاً. تؤكد هذه الحالة على الحاجة إلى استخراج بيانات منسقة من مستندات البروتوكول لتعزيز الكفاءة، وجودة الوثائق، والامتثال، مما يؤدي في النهاية إلى تحسين نتائج الرعاية الصحية.

تناقش الورقة قيود الطرق التقليدية لهيكلة البروتوكول وتشغيله، التي تعتمد بشكل كبير على المراجعات المدفوعة من الخبراء والتجريد اليدوي للبيانات. هذه الطرق ليست فقط تستغرق وقتًا طويلاً ولكنها أيضًا تقدم تناقضات وعدم دقة. يقترح المؤلفون الاستفادة من نماذج اللغة الكبيرة (LLMs) كحل تكنولوجي لأتمتة استخراج ورسم بيانات البروتوكول غير المنظمة. من خلال توليد نسخ أولية موحدة من البيانات والمستندات، يمكن لـ LLMs تبسيط سير العمل، مما يسمح للفرق بالتركيز على الأنشطة ذات الأهمية الأكبر مع تعزيز اتساق وقابلية تدقيق البيانات المستمدة من البروتوكول. تقيم الدراسة بشكل تجريبي نهجًا جديدًا لاستخراج معلومات التجارب السريرية، وتقارنه مع LLMs المستقلة عبر 23 بروتوكولًا وتقييم تأثيره التشغيلي من خلال تجربة محكومة مع منسقي الأبحاث السريرية (CRCs). تشير النتائج إلى تحسين الدقة في استخراج المعلومات المعقدة وتبرز التوفير الكبير في الوقت وتقليل العبء المعرفي على المستخدمين.

الطرق

توضح قسم “الطرق” الإجراءات التجريبية والتحليلية المستخدمة في الدراسة. استخدم الباحثون مزيجًا من الأساليب الكمية والنوعية لجمع البيانات، مما يضمن تحليلًا شاملاً للموضوع. تضمنت المنهجيات المحددة تجارب محكومة، واستطلاعات، وتحليلات إحصائية، تم تصميمها لاختبار الفرضيات التي تم وضعها في بداية البحث.

شمل جمع البيانات استخدام أدوات وبروتوكولات موحدة للحفاظ على الاتساق والموثوقية. تم اختيار عينة السكان بعناية لتمثيل التركيبة السكانية الأوسع ذات الصلة بالدراسة، وتم تطبيق تقنيات إحصائية مناسبة لتحليل النتائج. يبرز القسم أهمية الصرامة المنهجية في استخلاص استنتاجات صحيحة من النتائج، مما يسهم في قوة نتائج البحث.

النتائج

في التجربة المحكومة التي تقارن بين طرق استخراج البيانات المدعومة بالذكاء الاصطناعي واليدوية، تم تعيين 13 منسق أبحاث سريرية (CRCs) لبروتوكولين لكل منهم: واحد يستخدم استخراجًا مدعومًا بالذكاء الاصطناعي، والذي يتضمن جدول بيانات مسبق التعبئة للمراجعة والتصحيح، والآخر يستخدم عملية استخراج يدوية تبدأ من جدول بيانات فارغ. كانت الدراسة تهدف إلى استخراج البيانات من أربع فئات مختارة عشوائيًا من بين ست فئات محددة في القسم 3.2، مما أسفر عن إجمالي مقصود من 104 مهام استخراج. ومع ذلك، تضمنت مجموعة البيانات النهائية 97 مهمة فقط بسبب الاستخراج غير المكتمل للفئات من قبل أربعة منسقين واستبعاد المهام المتعلقة بجدول الأحداث المفقود لثلاثة منسقين.

توزيع المهام المكتملة عبر الفئات موضح في الجدول 4. أسفرت الطريقة المدعومة بالذكاء الاصطناعي عن 50 مهمة، بينما أسفرت الطريقة اليدوية عن 47 مهمة، مع إجمالي عدد المهام عبر الفئات البالغ 97. تضمنت الفئات الأحداث السلبية (19 مهمة)، المعلومات العامة (19 مهمة)، معايير الإدراج/الاستبعاد (12 مهمة)، التدخلات (21 مهمة)، المواقع (14 مهمة)، وجدول الأحداث (12 مهمة). تسلط هذه البيانات الضوء على الأداء المقارن لطريقتي الاستخراج من حيث إكمال المهام عبر فئات البيانات المختلفة.

المناقشة

تسلط قسم المناقشة في الورقة الضوء على قيود الأساليب الحالية التي تستخدم نماذج اللغة الكبيرة (LLMs) لاستخراج البيانات المنظمة من بروتوكولات التجارب السريرية، التي غالبًا ما تكون طويلة ومتنوعة. بينما تظهر الأدبيات الحالية استخدام LLMs لاستخراج ميزات محددة مثل معايير الأهلية والجداول، تقدم هذه الورقة خط أنابيب أكثر قوة لتوليد معزز بالاسترجاع (RAG) مصمم للتجارب السريرية. تتناول هذه الطريقة التحديات مثل قيود نافذة السياق، وتعقيد الجداول متعددة الصفحات، والحاجة إلى مخرجات دقيقة وقابلة للتدقيق. يؤكد المؤلفون على أهمية جدول الأحداث (SoE) في التنفيذ التشغيلي ويقترحون طريقة استخراج من مرحلتين تجمع بين اكتشاف الجداول مع قدرات LLM متعددة الوسائط لالتقاط العلاقات الهرمية في البيانات الجدولية بشكل فعال.

علاوة على ذلك، تناقش الورقة دمج LLMs في عملية التقييم، مما يوضح فعاليتها في تقييم جودة المحتوى الذي تم إنشاؤه بواسطة الذكاء الاصطناعي. يقدم المؤلفون منهجية هجينة للتعليق البشري-الذكاء الاصطناعي تعزز الكفاءة والدقة في استخراج البيانات مع الحفاظ على معايير عالية الجودة. التجربة المحكومة التي أجريت تقارن بين نهج الاستخراج المدعوم بالذكاء الاصطناعي والطرق اليدوية التقليدية، مما يكشف أن خط أنابيب RAG يحسن بشكل كبير من سرعة ودقة الاستخراج. بشكل عام، تشير النتائج إلى أن الإطار المقترح لـ RAG لا يعزز فقط عملية الاستخراج ولكن يوفر أيضًا حلاً قابلًا للتوسع لإدارة بيانات التجارب السريرية، مما يفيد أصحاب المصلحة المعنيين في جدوى المواقع، والعمليات، والمراقبة.

Journal: Journal of Biomedical Informatics, Volume: 179
DOI: https://doi.org/10.1016/j.jbi.2026.105036
PMID: https://pubmed.ncbi.nlm.nih.gov/41967790
Publication Date: 2026-04-10
Author(s): Ramtin Babaeipour et al.
Primary Topic: Scientific Computing and Data Management

Overview

The increasing complexity of clinical trial protocols, along with frequent amendments and challenges in knowledge management, poses significant burdens on trial teams. This study evaluates an Artificial Intelligence (AI) system that utilizes generative large language models (LLMs) combined with Retrieval-Augmented Generation (RAG) for the automated extraction of clinical trial protocol information. The research compares the extraction accuracy of the clinical-trial-specific RAG process, which achieved an accuracy of 89.0%, against standalone LLMs that yielded 62.6% accuracy when fine-tuned with prompts, using expert-supported reference annotations for validation.

Additionally, the operational impact of AI assistance on simulated workflows for Clinical Research Coordinators (CRCs) was assessed. Results indicated that AI-assisted tasks were completed at least 40% faster, were perceived as less cognitively demanding, and received strong user preference. While expert oversight remains crucial, these findings suggest that AI-assisted extraction can facilitate protocol intelligence at scale. This motivates the integration of such methodologies into real-world clinical workflows to further explore their effects on feasibility, study start-up, and post-activation monitoring.

Introduction

The introduction of this research paper highlights the critical role of clinical trial protocols in ensuring effective and safe medical interventions. It emphasizes that these protocols serve as essential communication tools among investigators, research teams, participants, and the scientific community, detailing the trial’s background, objectives, and design. However, the increasing complexity of protocols has led to variability in their completeness and adherence to quality guidelines, resulting in frequent, time-consuming amendments. This situation underscores the need for structured data extraction from protocol documents to enhance efficiency, documentation quality, and compliance, ultimately improving healthcare outcomes.

The paper discusses the limitations of traditional methods for protocol structuring and operationalization, which rely heavily on expert-driven reviews and manual data abstraction. These methods are not only time-consuming but also introduce inconsistencies and inaccuracies. The authors propose leveraging large language models (LLMs) as a technological solution to automate the extraction and mapping of unstructured protocol data. By generating standardized initial versions of data and documents, LLMs can streamline workflows, allowing teams to focus on higher-leverage activities while enhancing the consistency and auditability of protocol-derived data. The study empirically evaluates a novel clinical trial information extraction approach, comparing it against standalone LLMs across 23 protocols and assessing its operational impact through a controlled experiment with clinical research coordinators (CRCs). The findings indicate improved accuracy in extracting complex information and highlight significant time savings and reduced cognitive load for users.

Methods

The “Methods” section outlines the experimental and analytical procedures employed in the study. The researchers utilized a combination of quantitative and qualitative approaches to gather data, ensuring a comprehensive analysis of the subject matter. Specific methodologies included controlled experiments, surveys, and statistical analyses, which were designed to test the hypotheses formulated at the outset of the research.

Data collection involved the use of standardized instruments and protocols to maintain consistency and reliability. The sample population was carefully selected to represent the broader demographic relevant to the study, and appropriate statistical techniques were applied to analyze the results. The section emphasizes the importance of methodological rigor in drawing valid conclusions from the findings, ultimately contributing to the robustness of the research outcomes.

Results

In the controlled experiment comparing AI-assisted and manual data extraction methods, 13 Clinical Research Coordinators (CRCs) were assigned two protocols each: one utilizing AI-assisted extraction, which involved a pre-filled spreadsheet for review and correction, and the other employing a manual extraction process starting from an empty spreadsheet. The study aimed to extract data from four randomly selected categories out of six specified in Section 3.2, resulting in an intended total of 104 extraction tasks. However, the final dataset included only 97 tasks due to incomplete category extractions by four CRCs and the exclusion of tasks related to a missing schedule of events table for three CRCs.

The distribution of completed extraction tasks across categories is detailed in Table 4. The AI-assisted method yielded 50 tasks, while the manual method resulted in 47 tasks, with the total number of tasks across categories being 97. The categories included adverse events (19 tasks), general information (19 tasks), inclusion/exclusion criteria (12 tasks), interventions (21 tasks), sites (14 tasks), and schedule of events (12 tasks). This data highlights the comparative performance of the two extraction methods in terms of task completion across various data categories.

Discussion

The discussion section of the paper highlights the limitations of current approaches using large language models (LLMs) for extracting structured data from clinical trial protocols, which are often lengthy and heterogeneous. While existing literature demonstrates the use of LLMs for extracting specific features like eligibility criteria and schedules, this paper introduces a more robust Retrieval-Augmented Generation (RAG) pipeline tailored for clinical trials. This approach addresses challenges such as context window limitations, the complexity of multi-page tables, and the need for precise, auditable outputs. The authors emphasize the importance of the Schedule of Events (SoE) in operational execution and propose a two-stage extraction method that combines table detection with multimodal LLM capabilities to effectively capture hierarchical relationships in tabular data.

Furthermore, the paper discusses the integration of LLMs in the evaluation process, demonstrating their effectiveness in assessing the quality of AI-generated content. The authors present a hybrid human-AI annotation methodology that enhances efficiency and accuracy in data extraction while maintaining high-quality standards. The controlled experiment conducted compares the AI-assisted extraction approach with traditional manual methods, revealing that the RAG pipeline significantly improves extraction speed and accuracy. Overall, the findings suggest that the proposed RAG framework not only enhances the extraction process but also provides a scalable solution for clinical trial data management, benefiting stakeholders involved in site feasibility, operations, and monitoring.

شارك: