تمويل الذكاء الاصطناعي المسؤول ببيانات عالية الجودة من العالم الحقيقي: منصة S-RACE للبحث السريري القابل للتوسع ومتعدد التخصصات
Powering responsible artificial intelligence with high-quality real-world data: the S-RACE platform for scalable, multi-specialty clinical research

شارك:
المجلة: npj Digital Medicine، المجلد: 9، العدد: 1
DOI: https://doi.org/10.1038/s41746-025-02132-w
PMID: https://pubmed.ncbi.nlm.nih.gov/41484225
تاريخ النشر: 2026-01-03
المؤلف: Alberto Traverso وآخرون
الموضوع الرئيسي: الذكاء الاصطناعي في الرعاية الصحية والتعليم

نظرة عامة

تحدد هذه القسم تطوير وقدرات S-RACE، وهي منصة آمنة قائمة على السحابة تهدف إلى تحويل البيانات الصحية غير المنظمة إلى بيانات واقعية عالية الجودة (RWD) لتطبيقات الذكاء الاصطناعي السريرية. تستخدم S-RACE خط أنابيب شامل يبدأ بت anonymization محلي لضمان خصوصية البيانات، تليه تطبيق معالجة اللغة الطبيعية (NLP) لاستخراج وتوحيد المعلومات السريرية في تنسيق FHIR. تعتبر هذه البيانات المنظمة ضرورية لبناء نماذج ذكاء اصطناعي فعالة.

حاليًا، تحتوي S-RACE على بيانات من 31,276 مريضًا وتدعم 19 مشروع بحثي عبر مجالات طبية متنوعة، بما في ذلك الأورام، وأمراض القلب، والسكري. أثبتت المنصة فعاليتها من خلال التطبيقات في سرطان الكلى وتضيق الشريان الأورطي، حيث أظهرت نماذج الذكاء الاصطناعي المدربة على بيانات RWD المعالجة تلقائيًا من S-RACE مستويات أداء مقارنة بتلك المدربة على مجموعات بيانات تم تنسيقها يدويًا. من خلال توفير بيئة قابلة للتوسع ومحمية لتنسيق بيانات RWD، تؤسس S-RACE قاعدة موثوقة لتسهيل التكامل السريري لتقنيات الذكاء الاصطناعي المسؤولة.

مقدمة

تناقش مقدمة هذه الورقة البحثية تقاطع الأدلة الواقعية (RWE) والذكاء الاصطناعي (AI) في تحويل الطب الحديث، لا سيما من خلال تطوير أنظمة دعم القرار السريري المتقدمة (CDSS) التي تهدف إلى رعاية المرضى الشخصية. قدرة الذكاء الاصطناعي على تحليل مجموعات بيانات واسعة ومعقدة يمكن أن تكشف عن رؤى سريرية حاسمة، مما يحسن نتائج المرضى ويسرع اكتشاف الأدوية. ومع ذلك، فإن فعالية الذكاء الاصطناعي في الرعاية الصحية محدودة بشكل كبير بجودة بيانات العالم الحقيقي (RWD) التي يستخدمها، والتي غالبًا ما تكون غير منظمة، نادرة، وغير متسقة. تشكل هذه التباينات تحديات في توليد أدلة “ذات جودة تنظيمية” ضرورية للتطبيقات السريرية، حيث يمكن أن تؤدي جودة البيانات الضعيفة إلى نتائج غير موثوقة للذكاء الاصطناعي، بما في ذلك “الهلاوس” – النتائج التي لا تعكس بدقة البيانات الحقيقية.

لمعالجة هذه التحديات، تؤكد الورقة على الحاجة إلى منصات علم البيانات السريرية المتطورة القادرة على معالجة وتوحيد RWD بشكل منهجي لإنشاء مجموعات بيانات عالية الجودة جاهزة للتحليل. يبرز المؤلفون أهمية توافق عمليات إدارة البيانات مع المعايير التنظيمية المتطورة، مثل ISO/IEC 42001:2023 وقانون الذكاء الاصطناعي في الاتحاد الأوروبي، التي تفرض تدابير صارمة لجودة البيانات والحوكمة. استجابةً لهذه الاحتياجات، يقدم المؤلفون منصة S-RACE (San Raffaele Ai CEnter)، وهي حل قائم على السحابة مصمم لتعزيز جودة البيانات وحوكمتها في الذكاء الاصطناعي للرعاية الصحية. تتميز المنصة بخط أنابيب شامل لعلم البيانات يتضمن anonymization آمن للبيانات واستخراج البيانات المدفوع بـ NLP، مما يسهل في النهاية تحويل البيانات السريرية الخام إلى بيانات RWD ذات جودة بحثية. توضح الورقة بنية المنصة ووظائفها، مما يظهر دورها كبيئة تعاونية للأطباء وعلماء البيانات لتطوير أنظمة دعم القرار المدفوعة بالذكاء الاصطناعي المسؤولة، وبالتالي تحسين رعاية المرضى من خلال تحسين جودة البيانات وحوكمتها.

الطرق

تستفيد الطرق المستخدمة في هذا البحث من مجموعة من تقنيات Microsoft لإنشاء نظام بيئي شامل لمعالجة بيانات العالم الحقيقي (RWD) وتطوير الذكاء الاصطناعي (AI). تشمل المكونات الرئيسية خدمات الصحة المعرفية من Microsoft لمعالجة اللغة الطبيعية المتقدمة (NLP)، التي تستخرج معلومات طبية منظمة من نصوص سريرية غير منظمة باستخدام أنطولوجيات مثل نظام اللغة الطبية الموحد (UMLS). يتم تنظيم البيانات وفقًا لنموذج موارد التداخل السريع للرعاية الصحية (FHIR)، مع دعم إضافي لنموذج شراكة النتائج الطبية الملاحظة (OMOP) من خلال أدوات التحويل. يسهل Microsoft Power BI تصور البيانات بشكل بديهي، بينما يدعم Azure ML Studio دورة حياة التعلم الآلي (ML) بالكامل، مما يمكّن من تطوير نماذج متنوعة، بما في ذلك كل من النهج “الصندوق الأبيض” (مثل الانحدار اللوجستي) و”الصندوق الأسود” (مثل الغابات العشوائية)، مع أدوات تفسير مدمجة لتعزيز الفهم السريري.

تدمج المنصة البيانات من خمسة أنظمة تكنولوجيا المعلومات المؤسسية الرئيسية، بما في ذلك السجلات الصحية الإلكترونية (EHRs)، وأنظمة علم الأمراض والمختبرات، والتصوير الطبي (PACS)، وبيانات الجينوم، ومصادر البحث المحددة. تعزز هذه التكامل رؤية متعددة الأبعاد لبيانات المرضى، وهو أمر ضروري لـ 19 مشروع بحثي سريري جارٍ مفصل في الدراسة. يضمن نهج “الخصوصية حسب التصميم” الامتثال للوائح حماية البيانات العامة (GDPR) في الاتحاد الأوروبي من خلال إخفاء هوية بيانات المرضى، مع استبدال المعرفات المباشرة بتجزئات تشفيرية. تخضع معالجة البيانات لاتفاقية حماية البيانات الرسمية (DPA) مع Microsoft، مما يضمن الالتزام بمعايير GDPR. الطرق الموصوفة غير مرتبطة بالسحابة، مما يسمح بالتنفيذ عبر منصات متنوعة، وبالتالي تعزيز قابلية النقل وقابلية التوسع لإطار البحث.

النتائج

تستخدم منصة S-RACE نموذج حوكمة شامل لضمان توليد بيانات عالية الجودة، ذات جودة بحثية، تتماشى مع مبادئ الذكاء الاصطناعي المسؤول والأطر التنظيمية مثل قانون الذكاء الاصطناعي في الاتحاد الأوروبي وISO 42001:2023. ميزة رئيسية في هذا النموذج هي عملية تقييم جودة البيانات الهجينة التي تجمع بين تقييمات الخبراء مع سير العمل الآلي. تخضع مقترحات المشاريع لتمحيص صارم باستخدام قائمة مراجعة جودة البيانات، التي تقيم 39 معيارًا عبر خمس فئات، مما يسمح بتقييم منهجي لأبعاد جودة البيانات بما في ذلك الوصول، والدقة، والكمال، والتناسق، والملاءمة. بعد الموافقة على المشروع، يقوم إطار العمل الآلي لتحليل البيانات الاستكشافية الأولية (PExDA) بإجراء فحوصات جودة أساسية، مما يضمن أن البيانات “مناسبة للغرض” لأسئلة البحث المحددة.

تم بناء بنية المنصة على نهج “الخصوصية حسب التصميم”، مما يضمن الامتثال لـ GDPR من خلال عمليات مثل إخفاء هوية المعرفات المباشرة قبل نقل البيانات إلى السحابة. يسهل هذا الخط الأنبوبي الآمن للبيانات تحويل البيانات السريرية الخام إلى مجموعات بيانات عالية الجودة، منظمة وفقًا لمعيار FHIR. يدعم مختبر علم البيانات داخل المنصة تطوير نماذج التعلم الآلي، مع دمج تقنيات الذكاء الاصطناعي القابل للتفسير (XAI) لتعزيز شفافية النموذج وعدالته. من الجدير بالذكر أن المنصة قد أدارت بنجاح 19 مشروع بحثي سريري، مما يظهر قدرتها على توليد بيانات موثوقة لمجالات طبية متنوعة، بما في ذلك الأورام وأمراض القلب. يتم التحقق من فعالية منصة S-RACE بشكل أكبر من خلال المشاريع التي تستفيد من قدراتها في معالجة البيانات، مما يكشف أن نماذج الذكاء الاصطناعي المدربة على بيانات العالم الحقيقي المعالجة تلقائيًا تؤدي بشكل مشابه لتلك المستندة إلى مجموعات بيانات تم تنسيقها يدويًا، وبالتالي التغلب على حواجز كبيرة في تطوير الذكاء الاصطناعي القابل للتوسع للرعاية الصحية.

المناقشة

تم تطوير منصة S-RACE لتحويل البيانات المستشفى الخام وغير المتجانسة بشكل فعال إلى بيانات واقعية عالية الجودة، ذات جودة بحثية (RWD)، مما يعالج تحديًا حاسمًا في الذكاء الاصطناعي السريري. تؤكد على نهج صارم لدورة حياة البحث، مما يفعّل أطر مثل قائمة مراجعة PRINCIPLED لإعادة استخدام RWD. تشمل الميزات الرئيسية عملية تصميم دراسة منهجية تعتمد على إطار CRISP-DM، ودمج البيانات متعددة الأبعاد للتخفيف من التداخل، وبنية سحابية هجينة تعطي الأولوية لجودة البيانات والأمان من خلال إخفاء الهوية الإلزامي في الموقع. لا يعزز هذا التصميم الثقة بين أصحاب المصلحة فحسب، بل يتماشى أيضًا مع المعايير التنظيمية مثل ISO 42001:2023 وقانون الذكاء الاصطناعي في الاتحاد الأوروبي.

تتميز S-RACE عن المنصات الأخرى من خلال تعزيز التعاون بين الأطباء وعلماء البيانات من خلال بيئتين مزدوجتين: مركز الذكاء الاصطناعي للأطباء للمستخدمين غير التقنيين ومختبر علم البيانات للتحليلات المتقدمة. يدعم هذا الهيكل كل من التحليل المركزي والتعلم الفيدرالي الذي يحافظ على الخصوصية، مما يمكّن من تطوير نماذج ذكاء اصطناعي قابلة للتعميم مع الحفاظ على حوكمة بيانات صارمة. لقد أثبتت المنصة فعاليتها من خلال مشاريع سريرية متنوعة، مما يظهر قدرتها على تنسيق بيانات عالية الجودة وتسهيل تطبيقات الذكاء الاصطناعي المتقدمة. ومع ذلك، توجد قيود، لا سيما في مجال الاستدلال السببي والانحياز المحتمل في الاختيار بسبب ممارسات تنسيق البيانات. ستركز الأعمال المستقبلية على تعزيز قدرات الاستدلال السببي ومعالجة تحديات أداء النموذج بمرور الوقت.

Journal: npj Digital Medicine, Volume: 9, Issue: 1
DOI: https://doi.org/10.1038/s41746-025-02132-w
PMID: https://pubmed.ncbi.nlm.nih.gov/41484225
Publication Date: 2026-01-03
Author(s): Alberto Traverso et al.
Primary Topic: Artificial Intelligence in Healthcare and Education

Overview

The section outlines the development and capabilities of S-RACE, a secure, cloud-based platform aimed at transforming unstructured healthcare data into high-quality Real-World Data (RWD) for clinical AI applications. S-RACE employs an end-to-end pipeline that begins with on-premises anonymization to ensure data privacy, followed by the application of Natural Language Processing (NLP) to extract and standardize clinical information into the FHIR format. This structured data is crucial for constructing effective AI models.

Currently, S-RACE contains data from 31,276 patients and supports 19 research projects across various medical fields, including oncology, cardiology, and diabetes. The platform has demonstrated its effectiveness through applications in kidney cancer and aortic stenosis, where AI models trained on S-RACE’s automatically processed RWD exhibited performance levels comparable to those trained on manually curated datasets. By providing a scalable and governed environment for RWD curation, S-RACE establishes a reliable foundation to facilitate the clinical integration of responsible AI technologies.

Introduction

The introduction of this research paper discusses the intersection of Real-World Evidence (RWE) and Artificial Intelligence (AI) in transforming modern medicine, particularly through the development of advanced Clinical Decision Support Systems (CDSS) aimed at personalized patient care. AI’s ability to analyze extensive and complex datasets can uncover critical clinical insights, thereby improving patient outcomes and expediting drug discovery. However, the effectiveness of AI in healthcare is significantly limited by the quality of the Real-World Data (RWD) it utilizes, which is often unstructured, sparse, and inconsistent. This variability poses challenges in generating ‘regulatory grade’ evidence necessary for clinical applications, as poor data quality can lead to unreliable AI outcomes, including ‘hallucinations’—results that do not accurately reflect real data.

To address these challenges, the paper emphasizes the need for sophisticated clinical data science platforms capable of systematically processing and harmonizing RWD to create high-quality, analysis-ready datasets. The authors highlight the importance of aligning data management processes with evolving regulatory standards, such as ISO/IEC 42001:2023 and the EU AI Act, which mandate rigorous data quality and governance measures. In response to these needs, the authors introduce the S-RACE (San Raffaele Ai CEnter) platform, a cloud-based solution designed to enhance data quality and governance in healthcare AI. The platform features a comprehensive data science pipeline that includes secure data anonymization and NLP-driven data extraction, ultimately facilitating the transformation of raw clinical data into research-grade RWD. The paper outlines the platform’s architecture and functionalities, demonstrating its role as a collaborative environment for clinicians and data scientists to develop responsible AI-driven decision support systems, thereby enhancing patient care through improved data quality and governance.

Methods

The methods employed in this research leverage a suite of Microsoft technologies to create a comprehensive ecosystem for real-world data (RWD) processing and artificial intelligence (AI) development. Key components include Microsoft Cognitive Health Services for advanced Natural Language Processing (NLP), which extracts structured medical information from unstructured clinical texts using ontologies like the Unified Medical Language System (UMLS). The data is organized according to the Fast Healthcare Interoperability Resources (FHIR) model, with additional support for the Observational Medical Outcomes Partnership (OMOP) model through conversion tools. Microsoft Power BI facilitates intuitive data visualization, while Azure ML Studio supports the entire machine learning (ML) lifecycle, enabling the development of various models, including both “white box” (e.g., logistic regression) and “black box” (e.g., random forests) approaches, with integrated explainability tools to enhance clinical understanding.

The platform integrates data from five primary institutional IT systems, including Electronic Health Records (EHRs), pathology and laboratory systems, medical imaging (PACS), genomics data, and research-specific sources. This integration fosters a multimodal view of patient data, essential for the 19 ongoing clinical research projects detailed in the study. A ‘privacy by design’ approach ensures compliance with the EU’s General Data Protection Regulation (GDPR) through pseudonymisation of patient data, with direct identifiers replaced by cryptographic hashes. The data processing is governed by a formal Data Protection Addendum (DPA) with Microsoft, ensuring adherence to GDPR standards. The methods described are cloud-agnostic, allowing for implementation across various platforms, thus enhancing the portability and scalability of the research framework.

Results

The S-RACE platform employs a comprehensive governance model to ensure the generation of high-quality, research-grade data, adhering to Responsible AI principles and regulatory frameworks such as the EU AI Act and ISO 42001:2023. A key feature of this model is a hybrid data quality assessment process that combines expert evaluations with automated workflows. Project proposals undergo rigorous scrutiny using a Data Quality Checklist, which assesses 39 criteria across five categories, allowing for a systematic evaluation of data quality dimensions including Accessibility, Accuracy, Completeness, Consistency, and Relevancy. Following project approval, an automated Preliminary Exploratory Data Analysis (PExDA) framework conducts baseline quality checks, ensuring that the data is “fit for purpose” for specific research questions.

The platform’s architecture is built on a ‘privacy by design’ approach, ensuring compliance with GDPR through processes such as pseudonymisation of direct identifiers before data transfer to the cloud. This secure data pipeline facilitates the transformation of raw clinical data into high-quality datasets, structured according to the FHIR standard. The Data Science Lab within the platform supports the development of machine learning models, integrating Explainable AI (XAI) techniques to enhance model transparency and fairness. Notably, the platform has successfully managed 19 clinical research projects, demonstrating its capability to generate reliable data for diverse medical domains, including oncology and cardiovascular disease. The efficacy of the S-RACE platform is further validated through projects that leverage its data processing capabilities, revealing that AI models trained on automatically processed real-world data perform comparably to those based on manually curated datasets, thus overcoming significant barriers in scalable AI development for healthcare.

Discussion

The S-RACE platform has been developed to effectively transform raw, heterogeneous hospital data into high-quality, research-grade real-world data (RWD), addressing a critical challenge in clinical AI. It emphasizes a rigorous approach to the research lifecycle, operationalizing frameworks such as the PRINCIPLED checklist for RWD reuse. Key features include a systematic study design process based on the CRISP-DM framework, integration of multi-modal data to mitigate confounding, and a hybrid-cloud architecture that prioritizes data quality and security through mandatory on-premises pseudonymization. This design not only enhances trust among stakeholders but also aligns with regulatory standards like ISO 42001:2023 and the EU AI Act.

S-RACE distinguishes itself from other platforms by fostering collaboration between clinicians and data scientists through dual environments: the Clinician AI Hub for non-technical users and the Data Science Lab for advanced analytics. This structure supports both centralized analysis and privacy-preserving federated learning, enabling the development of generalizable AI models while maintaining stringent data governance. The platform has demonstrated its effectiveness through various clinical projects, showcasing its ability to curate high-quality data and facilitate advanced AI applications. However, limitations exist, particularly in the realm of causal inference and potential selection bias due to data curation practices. Future work will focus on enhancing causal inference capabilities and addressing the challenges of model performance over time.

شارك: