تمكين إدارة البيانات العادلة في الدراسات الدولية متعددة المواقع المعقدة: عمليات البيانات لشراكة تسريع الأدوية® برنامج الفصام
Enabling FAIR data stewardship in complex international multi-site studies: Data Operations for the Accelerating Medicines Partnership® Schizophrenia Program

شارك:
المجلة: Schizophrenia، المجلد: 11، العدد: 1
DOI: https://doi.org/10.1038/s41537-025-00560-x
PMID: https://pubmed.ncbi.nlm.nih.gov/40180978
تاريخ النشر: 2025-04-03
المؤلف: Tashrif Billah وآخرون
الموضوع الرئيسي: الحوسبة العلمية وإدارة البيانات

نظرة عامة

تحدد هذه القسم نموذج حوكمة البيانات المعتمد لمشروع شراكة تسريع الأدوية® للفصام، مع التأكيد على دور المستفيدين كأوصياء على البيانات بدلاً من مالكين. يتطلب هذا النموذج مشاركة واسعة للبيانات المجمعة مع التنقل عبر تعقيدات التعامل مع المعلومات الصحية الحساسة، خاصة في الدراسات التي تشمل الأفراد الشباب المعرضين لخطر عالٍ للإصابة بالاضطرابات النفسية. يوضح المؤلفون نظام تشغيل البيانات، الذي يشمل أنظمة لجمع البيانات، وتنظيمها، وتدفقها، وأمانها، ومراقبة الجودة، والتصور، والمراقبة، والنشر عبر منصة أرشيف بيانات المعهد الوطني للصحة النفسية.

تركز المشروع بشكل رئيسي على دمج مبادئ FAIR (قابلية الاكتشاف، الوصول، التشغيل البيني، وإعادة الاستخدام) في عملية إدارة البيانات، مما يضمن الامتثال للمعايير التنظيمية والأخلاقية المحلية. تعزز تنفيذ خطوط الأنابيب الآلية جودة البيانات والتعاون، بينما يسمح التركيز على تقييم الجودة (QA) ومراقبة الجودة (QC) القريب من الوقت الحقيقي بإجراء تصحيحات وتحسينات فورية في جمع البيانات. لا يدعم هذا النهج نزاهة البحث فحسب، بل يعزز أيضًا المنهجيات للدراسات متعددة المواقع التي تشمل حالات الصحة النفسية الحساسة.

مقدمة

يهدف مشروع شراكة تسريع الأدوية للفصام (AMP SCZ) الرصدي إلى تعزيز فهم متلازمة الخطر العالي السريري (CHR) وتحسين التنبؤ ببدء نوبات الفصام. نظرًا للاختلاف الكبير في الأعراض والنتائج، يسعى المبادرة إلى تطوير أساليب الطب الدقيق من خلال استخدام مجموعات بيانات كبيرة. يعتزم المشروع متابعة 2,617 مشاركًا تتراوح أعمارهم بين 12-30 عامًا على مدى عامين، وجمع بيانات شاملة عبر مجالات متنوعة، بما في ذلك التقييمات السريرية، وتخطيط الدماغ الكهربائي (EEG)، والتصوير بالرنين المغناطيسي (MRI)، والمعرفة العصبية، والبيانات الحيوية السائلة، وتقنيات الصحة الرقمية، والتسجيلات السمعية البصرية.

تلتزم AMP SCZ بممارسات العلوم المفتوحة، مما يضمن مشاركة بيانات المستوى الشخصي غير المحددة الهوية مع المجتمع البحثي الأوسع. تتطلب تعقيدات تنسيق إدارة البيانات عبر أكثر من 150 عالمًا من خمس قارات بنية تحتية قوية لجمع البيانات، ومعالجتها، وضمان الجودة، والنشر، وهو أمر حاسم للحفاظ على جودة البيانات العالية ودعم الأهداف العلمية للمشروع.

الطرق

تحدد قسم “الطرق” في ورقة البحث التصميم التجريبي والتقنيات التحليلية المستخدمة للتحقيق في أسئلة البحث. استخدمت الدراسة نهجًا كميًا، حيث تم دمج التحليلات الإحصائية لتقييم البيانات المجمعة من تجارب متنوعة. شملت المنهجيات المحددة تجارب مختبرية محكومة، حيث تم التلاعب بالمتغيرات بشكل منهجي لملاحظة تأثيراتها على النتائج المعنية.

شمل جمع البيانات استخدام أدوات موحدة لضمان الموثوقية والصلاحية. تم إجراء التحليل باستخدام برامج إحصائية مناسبة، مع تطبيق اختبارات مثل ANOVA وتحليل الانحدار لتحديد الفروق والعلاقات المهمة بين المتغيرات. يبرز القسم أهمية القابلية للتكرار والشفافية في الطرق المستخدمة، موضحًا حجم العينة، ومعايير الاختيار، وأي تحيزات محتملة تم التعامل معها خلال الدراسة. بشكل عام، يدعم الصرامة المنهجية التي تم تأسيسها في هذا القسم مصداقية النتائج المقدمة في الورقة.

النتائج

تسلط قسم النتائج الضوء على الأبعاد الحرجة للحجم، والسرعة، والتنوع في سياق مبادرة AMP SCZ، وهي دراسة جماعية مستقبلية واسعة النطاق تركز على الأفراد المعرضين لخطر عالٍ سريريًا (CHR) للإصابة بالفصام. تمثل الدراسة حجم بيانات كبير، حيث يمكن أن يساهم كل مشارك بما يصل إلى 11,324 متغيرًا، بما في ذلك بيانات عصبية فسيولوجية وتصوير عصبي واسعة النطاق، بالإضافة إلى جمع بيانات مستمرة من الهواتف الذكية وأجهزة قياس التسارع. يمكن أن تصل البيانات الخام المجمعة لمشارك واحد إلى حوالي 50 جيجابايت، واعتبارًا من فبراير 2024، جمعت المبادرة ما يقرب من 15 تيرابايت من البيانات الخام، مع تقديرات تشير إلى أكثر من 100 تيرابايت عند الانتهاء من الدراسة التي تشمل حوالي 2000 فرد من CHR و640 من ضوابط المجتمع المتطابقة.

يؤكد القسم أيضًا على سرعة توليد البيانات، مما يتطلب معالجة قريبة من الوقت الحقيقي لاستخراج رؤى في الوقت المناسب، وتنوع أنواع البيانات، التي تتراوح من التنسيقات المنظمة إلى غير المنظمة. لا توضح هذه الأبعاد فقط التحديات المرتبطة بإدارة وتحليل مثل هذه المجموعات الكبيرة من البيانات، بل تبرز أيضًا الفرص لاستخدام البيانات بشكل مبتكر عبر مختلف التخصصات. تتطلب البنية التحتية اللازمة لدعم هذا الجمع الواسع للبيانات موارد بشرية ووقت كبيرين، مما يبرز تعقيد الحفاظ على المشروع طوال مدته.

المناقشة

تتناول قسم المناقشة في ورقة البحث دمج DataOps ومبادئ FAIR ضمن برنامج AMP SCZ، مع التأكيد على الحاجة إلى تحسين إدارة البيانات في سياق مشاريع البيانات واسعة النطاق. يتميز DataOps بتركيزه على تعزيز دورة حياة البيانات من خلال تدفق البيانات السلس، ودورات التغذية الراجعة السريعة، والتحسين المستمر للعملية. يعطي برنامج AMP SCZ الأولوية لتقديم بيانات عالية الجودة للمستفيدين ويستخدم إطار حوكمة قوي للإشراف على نزاهة البيانات والتعامل الأخلاقي عبر شبكته الواسعة المكونة من 43 موقعًا دوليًا. هذه المبادرة غير مسبوقة في حجمها وتعقيدها، مما يتطلب تنسيقًا دقيقًا بين فرق متنوعة وأنظمة جمع البيانات.

علاوة على ذلك، تبرز الورقة الالتزام بمبادئ FAIR—قابلية الاكتشاف، الوصول، التشغيل البيني، وإعادة الاستخدام—كأساسية لتعظيم تأثير العلوم المفتوحة. يهدف البرنامج إلى تسهيل مشاركة البيانات من خلال أرشيف بيانات المعهد الوطني للصحة النفسية مع ضمان خصوصية المشاركين من خلال تدابير أخلاقية صارمة. يبرز تنفيذ أنظمة جمع البيانات الإلكترونية المختلفة، بما في ذلك REDCap وRPMS، جنبًا إلى جنب مع أدوات تجميع البيانات المبتكرة مثل Lochness، التزام البرنامج بالحفاظ على معايير عالية من جودة البيانات والأمان. بشكل عام، يمثل برنامج AMP SCZ نهجًا شاملاً لعمليات البيانات التي تعزز كفاءة البحث وتعزز الاستكشاف العلمي التعاوني.

Journal: Schizophrenia, Volume: 11, Issue: 1
DOI: https://doi.org/10.1038/s41537-025-00560-x
PMID: https://pubmed.ncbi.nlm.nih.gov/40180978
Publication Date: 2025-04-03
Author(s): Tashrif Billah et al.
Primary Topic: Scientific Computing and Data Management

Overview

The section outlines the data governance model adopted for the Accelerating Medicines Partnership® Schizophrenia project, emphasizing the role of grantees as stewards of data rather than owners. This model necessitates the broad sharing of collected data while navigating the complexities of handling sensitive health information, particularly in studies involving young individuals at clinical high risk for psychotic disorders. The authors detail the data operation ecosystem, which includes systems for data capture, organization, flow, security, quality control, visualization, monitoring, and dissemination via the NIMH Data Archive platform.

A key focus of the project is the integration of the FAIR (Findability, Accessibility, Interoperability, and Reusability) principles into the data management process, ensuring compliance with local regulatory and ethical standards. The implementation of automated pipelines enhances data quality and collaboration, while the emphasis on close-to-real-time quality assessment (QA) and quality control (QC) allows for immediate corrections and improvements in data acquisition. This approach not only supports the integrity of the research but also advances methodologies for multisite studies involving sensitive mental health conditions.

Introduction

The Accelerating Medicines Partnership Schizophrenia (AMP SCZ) observational project aims to enhance the understanding of clinical high-risk (CHR) syndrome and improve the prediction of psychotic episode onset. Given the significant heterogeneity in symptoms and outcomes, the initiative seeks to develop precision medicine approaches by utilizing large datasets. The project intends to follow 2,617 participants aged 12-30 over two years, gathering comprehensive data across various domains, including clinical assessments, electroencephalography (EEG), magnetic resonance imaging (MRI), neurocognition, fluid biomarkers, digital health technologies, and audiovisual recordings.

AMP SCZ is committed to open science practices, ensuring that de-identified subject-level data will be shared with the broader research community. The complexity of coordinating data management across over 150 scientists from five continents necessitates a robust infrastructure for data capture, processing, quality assurance, and dissemination, which is crucial for maintaining high data quality and supporting the project’s scientific goals.

Methods

The “Methods” section of the research paper outlines the experimental design and analytical techniques employed to investigate the research questions. The study utilized a quantitative approach, incorporating statistical analyses to evaluate the data collected from various experiments. Specific methodologies included controlled laboratory experiments, where variables were systematically manipulated to observe their effects on the outcomes of interest.

Data collection involved the use of standardized instruments to ensure reliability and validity. The analysis was conducted using appropriate statistical software, with tests such as ANOVA and regression analysis applied to determine significant differences and relationships among the variables. The section emphasizes the importance of replicability and transparency in the methods used, detailing the sample size, selection criteria, and any potential biases that were addressed during the study. Overall, the methodological rigor established in this section supports the credibility of the findings presented in the paper.

Results

The results section highlights the critical dimensions of volume, velocity, and variety in the context of the AMP SCZ initiative, a large-scale prospective cohort study focused on individuals at clinical high risk (CHR) for psychosis. The study exemplifies significant data volume, with each participant potentially contributing up to 11,324 variables, including extensive neurophysiological and neuroimaging data, as well as continuous smartphone and accelerometry data collection. The cumulative raw data collected for a single participant can reach approximately 50 GB, and as of February 2024, the initiative has amassed nearly 15 TB of raw data, with projections estimating over 100 TB upon completion of the study involving around 2000 CHR individuals and 640 matched community controls.

The section further emphasizes the velocity of data generation, necessitating near-real-time processing to derive timely insights, and the variety of data types, which range from structured to unstructured formats. These dimensions not only illustrate the challenges associated with managing and analyzing such vast datasets but also highlight the opportunities for innovative data utilization across various disciplines. The infrastructure required to support this extensive data collection involves significant human resources and time, underscoring the complexity of maintaining the project throughout its duration.

Discussion

The discussion section of the research paper elaborates on the integration of DataOps and FAIR principles within the AMP SCZ program, emphasizing the need for optimized data management in the context of large-scale data projects. DataOps is characterized by its focus on enhancing the data lifecycle through seamless data flow, rapid feedback loops, and continuous process improvement. The AMP SCZ program prioritizes high-quality data delivery to stakeholders and employs a robust governance framework to oversee data integrity and ethical handling across its extensive network of 43 international sites. This initiative is unprecedented in its scale and complexity, necessitating meticulous coordination among diverse teams and data capture systems.

Furthermore, the paper highlights the adherence to FAIR principles—Findability, Accessibility, Interoperability, and Reusability—as essential for maximizing the impact of open science. The program aims to facilitate data sharing through the NIMH Data Archive while ensuring participant privacy through stringent ethical safeguards. The implementation of various electronic data capture systems, including REDCap and RPMS, alongside innovative data aggregation tools like Lochness, underscores the program’s commitment to maintaining high standards of data quality and security. Overall, the AMP SCZ program exemplifies a comprehensive approach to data operations that not only enhances research efficiency but also fosters collaborative scientific exploration.

شارك: