دمج متعدد الأوميك لبيانات الميكروبيوم لتحديد الوحدات المرتبطة بالمرض
Multi-omic integration of microbiome data for identifying disease-associated modules

المجلة: Nature Communications، المجلد: 15، العدد: 1
DOI: https://doi.org/10.1038/s41467-024-46888-3
PMID: https://pubmed.ncbi.nlm.nih.gov/38521774
تاريخ النشر: 2024-03-23
المؤلف: Efrat Muller وآخرون
الموضوع الرئيسي: الميكروبيوم المعوي والصحة

طرق

قسم “الطرق” يوضح الأساليب التجريبية والتحليلية المستخدمة في الدراسة. يوضح معايير اختيار المشاركين، وتصميم التجارب، والتقنيات الإحصائية المستخدمة لتحليل البيانات. يتم وصف منهجيات محددة، مثل التجارب المضبوطة أو الدراسات الملاحظة، بالإضافة إلى أي أدوات أو أجهزة مستخدمة للقياس.

بالإضافة إلى ذلك، قد يتضمن القسم معلومات عن حجم العينة، وإجراءات جمع البيانات، وأي اعتبارات أخلاقية تم أخذها في الاعتبار خلال البحث. تم تصميم الطرق لضمان موثوقية وصلاحية النتائج، مما يسمح بإعادة الإنتاج والتعميم للنتائج ضمن المجال ذي الصلة بالدراسة.

نتائج

يقدم القسم نتائج تطوير MintTea، وهو إطار تكامل متعدد الأوميات مصمم لتحديد وحدات الميكروبيوم المرتبطة بالأمراض. يستخدم MintTea تحليل الارتباط العام المتناثر (sGCCA) جنبًا إلى جنب مع أخذ عينات متكررة وتحليل الإجماع لتقليل الضوضاء وتعزيز موثوقية النتائج. يعالج الإطار جداول ميزات متعددة الأوميات وعلامة مرض لكل عينة، مع تصفية الميزات النادرة وترميز العلامة كأوم إضافي. من خلال التطبيق التكراري لـ sGCCA، ينتج MintTea متغيرات كامنة تمثل تركيبات خطية متناثرة من الميزات عبر أوميات مختلفة، مما يؤدي إلى تحديد “وحدات افتراضية”.

لضمان قوة هذه الوحدات، يقوم MintTea بإجراء عدة تكرارات على مجموعات فرعية عشوائية من البيانات، ويقوم ببناء شبكة تزامن لتحديد “وحدات الإجماع” بناءً على تزامن الميزات المتسق عبر التكرارات. يوازن الإطار بين أهداف تعظيم الارتباطات بين الأوميات وارتباطها بالمرض. يتم تقييم وحدات الإجماع النهائية بناءً على متوسط الارتباط عبر الأوميات وقوة التنبؤ، التي يتم قياسها باستخدام المساحة تحت منحنى التشغيل (AUC). يتم الاحتفاظ فقط بالوحدات التي لديها AUC أكبر من 0.7 وأفضل من الوحدات العشوائية، مما يؤدي إلى عدة وحدات متعددة الأوميات مرتبطة بالأمراض التي تلتقط ميزات مرتبطة بشدة عبر أوميات مختلفة وعلاقتها بالمرض. يختتم القسم بتسليط الضوء على مزايا MintTea مقارنة بأساليب أخرى تعتمد على sGCCA، وخاصة معدل الاكتشاف الخاطئ المنخفض والتوازن الفعال للخصائص المرغوبة.

مناقشة

في هذه الدراسة، قدم المؤلفون MintTea، وهو إطار مصمم لتحديد وحدات متعددة الأوميات المرتبطة بالأمراض من مجموعات بيانات الميكروبيوم المتنوعة. باستخدام تسع مجموعات بيانات حالة-تحكم تشمل حالات صحية مختلفة، بما في ذلك مرض كرون، والتهاب القولون التقرحي، ومتلازمة الأيض، أنتج MintTea عدة جداول ميزات بناءً على الملفات التصنيفية والوظيفية من بيانات الميتاجينوميات والميتابولوميات. كشفت التحليلات عن 2-5 وحدات إجماع لكل مجموعة بيانات، مع ظهور العديد من الوحدات التي تظهر ارتباطات كبيرة بين الميزات عبر أوميات مختلفة، متجاوزة تلك الخاصة بالوحدات المأخوذة عشوائيًا. ومن الجدير بالذكر أن بعض الوحدات أظهرت ارتباطات قوية مع أمراض معينة، محققة قيم المساحة تحت المنحنى (AUC) التي تشير إلى قوة تنبؤية قوية.

سلطت النتائج الضوء على إمكانيات طرق التكامل الوسيطة مثل MintTea لكشف التفاعلات البيولوجية المعقدة ضمن بيانات متعددة الأوميات. على سبيل المثال، تضمنت الوحدات المرتبطة بسرطان القولون والمستقيم أنواعًا معينة من البكتيريا والأحماض الأمينية، بينما كشفت الوحدات المرتبطة بمتلازمة الأيض عن تجمعات من المستقلبات المصلية وأنواع البكتيريا المرتبطة بالاضطرابات الأيضية. كما حددت الدراسة وحدات متعددة الأوميات المتكررة عبر مجموعات البيانات، مما يشير إلى توقيعات بيولوجية شائعة تتشاركها مجموعات مختلفة. وأكد المؤلفون على أهمية تقييم قوة وعمومية وملاءمة الوحدات البيولوجية المحددة، مع الاعتراف بالتحديات في التحقق من التفاعلات متعددة الأوميات بسبب توفر البيانات وتعقيد الأنظمة البيولوجية. بشكل عام، يعد MintTea أداة قيمة لتوليد فرضيات حول دور الميكروبيوم في الأمراض، مما يمهد الطريق للبحوث المستقبلية في تكامل الأوميات.

Journal: Nature Communications, Volume: 15, Issue: 1
DOI: https://doi.org/10.1038/s41467-024-46888-3
PMID: https://pubmed.ncbi.nlm.nih.gov/38521774
Publication Date: 2024-03-23
Author(s): Efrat Muller et al.
Primary Topic: Gut microbiota and health

Methods

The “Methods” section outlines the experimental and analytical approaches employed in the study. It details the selection criteria for participants, the design of the experiments, and the statistical techniques used for data analysis. Specific methodologies, such as controlled trials or observational studies, are described, along with any tools or instruments utilized for measurement.

Additionally, the section may include information on the sample size, data collection procedures, and any ethical considerations taken into account during the research. The methods are designed to ensure the reliability and validity of the findings, allowing for reproducibility and generalization of results within the relevant field of study.

Results

The section presents the results of the development of MintTea, a multi-omic integration framework designed to identify robust disease-associated microbiome modules. MintTea employs sparse generalized canonical correlation analysis (sGCCA) alongside repeated sampling and consensus analysis to mitigate noise and enhance the reliability of results. The framework processes multiple omic feature tables and a disease label for each sample, filtering rare features and encoding the label as an additional omic. Through iterative application of sGCCA, MintTea generates latent variables that represent sparse linear combinations of features across different omics, leading to the identification of “putative modules.”

To ensure the robustness of these modules, MintTea conducts multiple iterations on random subsets of the data, constructing a co-occurrence network to identify “consensus modules” based on consistent feature co-occurrence across iterations. The framework balances the objectives of maximizing correlations between omics and their association with the disease. Final consensus modules are evaluated based on their average cross-omic correlation and predictive power, quantified using the area under the receiver operating characteristic curve (AUC). Only modules with an AUC greater than 0.7 and superior to random modules are retained, resulting in multiple disease-associated, multi-omic modules that capture highly correlated features across different omics and their relationship with the disease. The section concludes by highlighting MintTea’s advantages over other sGCCA-based methods, particularly its lower false discovery rate and effective balance of desired properties.

Discussion

In this study, the authors introduced MintTea, a framework designed to identify disease-associated multi-omic modules from diverse microbiome datasets. Utilizing nine case-control datasets encompassing various health conditions, including Crohn’s disease, ulcerative colitis, and metabolic syndrome, MintTea generated multiple feature tables based on taxonomic and functional profiles from metagenomics and metabolomics data. The analysis revealed 2-5 consensus modules per dataset, with many modules demonstrating significant correlations between features across different omics, surpassing those of randomly sampled modules. Notably, some modules exhibited strong associations with specific diseases, achieving area under the curve (AUC) values that indicated robust predictive power.

The findings highlighted the potential of intermediate integration methods like MintTea to uncover complex biological interactions within multi-omic data. For instance, modules associated with colorectal cancer included specific bacterial species and amino acids, while those linked to metabolic syndrome revealed clusters of serum metabolites and bacterial species implicated in metabolic disorders. The study also identified recurring multi-omic modules across datasets, suggesting common biological signatures shared among different cohorts. The authors emphasized the importance of evaluating the robustness, generalizability, and biological relevance of the identified modules, acknowledging the challenges in validating multi-omic interactions due to data availability and the complexity of biological systems. Overall, MintTea serves as a valuable tool for generating hypotheses regarding the microbiome’s role in disease, paving the way for future research in multi-omic integration.