تجميع عالي الجودة للميتا جينوم من قراءات النانو بورا باستخدام nanoMDBG
High-quality metagenome assembly from nanopore reads with nanoMDBG

شارك:
المجلة: Nature Communications، المجلد: 17، العدد: 1
DOI: https://doi.org/10.1038/s41467-026-69760-y
PMID: https://pubmed.ncbi.nlm.nih.gov/41792155
تاريخ النشر: 2026-03-06
المؤلف: Gaëtan Benoit وآخرون
الموضوع الرئيسي: دراسات الجينوميات والتطور

نظرة عامة

يتناول هذا القسم التقدم في تجميع الميتاجينوم الذي يسهل استخدام تقنيات التسلسل الطويلة من الجيل الثالث، مع تسليط الضوء بشكل خاص على قدرات PacBio HiFi وتقنيات أكسفورد نانوبور (ONT). تُعرف قراءات PacBio HiFi بدقتها العالية، مما يمكن من إنتاج مئات من الجينومات المجمعة من الميتاجينوم شبه الكاملة (MAGs) من عينة واحدة. في الوقت نفسه، أدت التحسينات الأخيرة في ONT إلى تقليل معدل الخطأ لكل قاعدة إلى 1-2%، ومع ذلك، فإن مجمعات الميتاجينوم الحالية مُحسّنة بشكل أساسي لقراءات HiFi وتواجه صعوبة مع مجموعات البيانات الأكبر التي تنتجها ONT.

لمعالجة هذه القيود، يقدم المؤلفون nanoMDBG، وهو إصدار محسّن من metaMDBG يتضمن خطوة معالجة مسبقة لتصحيح الأخطاء في مساحة المصغرات، مصممة خصيصًا لقراءات ONT. تم عرض أداء nanoMDBG عبر مجموعات بيانات ONT المختلفة، بما في ذلك عينة تربة كبيرة بحجم 400 جيجابايت، حيث نجح في إعادة بناء ما يصل إلى ضعف عدد MAGs عالية الجودة مقارنةً بأفضل مجمع ONT، metaFlye، بينما استخدم ثلث الوقت والذاكرة فقط. ومن الجدير بالذكر أن النتائج تشير إلى أن أحدث تقنيات ONT يمكن أن تحقق نتائج بناء MAG مقارنة بتلك الخاصة بـ PacBio HiFi عند أعماق تسلسل متساوية.

مقدمة

تناقش مقدمة ورقة البحث التقدم في الميتاجينوميات، مع التركيز بشكل خاص على تجميع قراءات التسلسل من عينات بيئية لتحليل المجتمعات الميكروبية المعقدة. غالبًا ما تؤدي طرق التسلسل القصير التقليدية إلى تجميعات مجزأة وجينومات ميتاجينوم مجمعة غير مكتملة (MAGs)، مما يزيد من تعقيدها بسبب مشكلات التلوث والتجزئة العالية. بالمقابل، أدت تقنيات التسلسل الطويل، مثل تلك التي تقدمها أكسفورد نانوبور (ONT) وPacBio HiFi، إلى تحسين كبير في جودة التجميع. تمكنت قراءات PacBio HiFi، المعروفة بدقتها العالية، من حل العديد من الجينومات الدائرية من عينات معقدة، بينما أدت التحسينات الأخيرة في تقنية ONT إلى تقليل معدلات الخطأ إلى حوالي 1%، مما يسهل تطوير منهجيات جديدة.

تسلط المقدمة الضوء على العديد من المجمعات الرئيسية التي تم تطويرها لبيانات الميتاجينوم الطويلة، بما في ذلك MetaFlye، التي كانت الأولى المصممة خصيصًا لهذا الغرض، وHifiasm-meta، التي تتفوق في إعادة بناء الجينومات الكاملة والدائرية باستخدام قراءات HiFi. علاوة على ذلك، تذكر المقدمة التطور الأخير لـ metaMDBG، الذي يُحسن تجميع بيانات HiFi من خلال تقليل استخدام الذاكرة بشكل كبير وتجاوز الأدوات الحالية في إنتاج MAGs شبه الكاملة وكفاءة وقت التشغيل. تؤكد هذه التقدمات على إمكانيات تقنيات التسلسل الطويل لتعزيز فهمنا للتنوع والوظائف الميكروبية في بيئات مختلفة.

طرق

في قسم الطرق، يحدد المؤلفون الإطار الإجرائي لنهج nanoMDBG، كما هو موضح في الشكل 1. يشمل هذا الإطار سلسلة من الخطوات المنهجية المصممة لتسهيل تنفيذ nanoMDBG في أبحاثهم. يتم توضيح كل خطوة في الأقسام اللاحقة، مما يوفر وضوحًا حول المنهجيات المستخدمة والأسباب وراءها. تهدف الوصف التفصيلي إلى ضمان إمكانية إعادة الإنتاج وتسليط الضوء على الجوانب الابتكارية لتقنية nanoMDBG.

نتائج

في هذا القسم، يقدم المؤلفون نتائج تقييماتهم للتجميع، بدءًا من تحليل التجميعات المشتقة فقط من بيانات أكسفورد نانوبور (ONT). بعد ذلك، يقارنون التجميعات المنتجة باستخدام كل من بيانات ONT وبيانات عالية الدقة (HiFi) من عينات متطابقة. ينتهي القسم بتقييم أخطاء التجميع التي لوحظت عبر كلا منصتي التسلسل، مع تسليط الضوء على الفروق في الأداء والدقة بين المنهجين. تهدف هذه التقييم الشامل إلى توضيح نقاط القوة والضعف لكل نوع بيانات في عمليات تجميع الجينوم.

نقاش

في هذا القسم، يقدم المؤلفون nanoMDBG، وهو طريقة جديدة لتجميع الميتاجينوم مصممة لبيانات التسلسل الطويلة من أكسفورد نانوبور (ONT)، والتي تعزز المجمع الحالي metaMDBG. تستخدم الطريقة نهج مساحة المصغرات لتصحيح القراءات وتجميع الميتاجينومات، مما يحسن بشكل كبير من إعادة بناء الجينومات المجمعة من الميتاجينوم (MAGs)، الفيروسات، والبلازميدات. يكشف التقييم مقارنةً بالمجمعات المتطورة، بما في ذلك metaFlye وmetaMDBG، أن nanoMDBG يتفوق باستمرار على هذه الأدوات من حيث عدد MAGs شبه الكاملة، خاصة في مجموعات البيانات المعقدة. على سبيل المثال، في عينات الأمعاء البشرية والتربة، أعاد nanoMDBG بناء عدد أكبر بكثير من MAGs شبه الكاملة وMAGs ذات القطعة الواحدة مقارنةً بنظرائه.

كما يبرز المؤلفون أنه بينما تظهر بيانات ONT وعدًا، خاصة من حيث التكلفة والقدرة على التوسع لدراسات الميتاجينوم، لا تزال هناك تحديات في تحقيق تغطية عالية للمجتمعات المعقدة مثل التربة. تشير النتائج إلى أن 37% فقط من الجينومات عالية التغطية تم تحويلها إلى MAGs من مجموعة بيانات التربة، مقارنةً بـ 85% من الأمعاء البشرية، مما يشير إلى مجالات محتملة لمزيد من التحسين. بشكل عام، يمثل nanoMDBG تقدمًا كبيرًا في تجميع الميتاجينوم، مما يمكّن من تحليلات أكثر شمولية وفعالية من حيث التكلفة للمجتمعات الميكروبية، مع آثار على تطبيقات مثل تتبع السلالات وفهم الديناميات التطورية.

Journal: Nature Communications, Volume: 17, Issue: 1
DOI: https://doi.org/10.1038/s41467-026-69760-y
PMID: https://pubmed.ncbi.nlm.nih.gov/41792155
Publication Date: 2026-03-06
Author(s): Gaëtan Benoit et al.
Primary Topic: Genomics and Phylogenetic Studies

Overview

The section discusses advancements in metagenome assembly facilitated by third-generation long-read sequencing technologies, particularly highlighting the capabilities of PacBio HiFi and Oxford Nanopore Technologies (ONT). PacBio HiFi reads are noted for their high accuracy, enabling the generation of hundreds of near-complete metagenome-assembled genomes (MAGs) from a single sample. Meanwhile, ONT’s recent improvements have reduced its per-base error rate to 1-2%, yet existing metagenome assemblers are primarily optimized for HiFi reads and struggle with the larger datasets produced by ONT.

To address this limitation, the authors introduce nanoMDBG, an enhanced version of metaMDBG that incorporates an error correction pre-processing step in minimizer-space, specifically designed for ONT reads. The performance of nanoMDBG is demonstrated across various ONT datasets, including a substantial 400 Gbp soil sample, where it successfully reconstructs up to twice as many high-quality MAGs compared to the leading ONT assembler, metaFlye, while utilizing only one-third of the CPU time and memory. Notably, the results indicate that the latest ONT technology can achieve MAG construction outcomes comparable to those of PacBio HiFi at equivalent sequencing depths.

Introduction

The introduction of the research paper discusses the advancements in metagenomics, particularly focusing on the assembly of sequencing reads from environmental samples to analyze complex microbial communities. Traditional short-read sequencing methods often result in fragmented assemblies and incomplete metagenome-assembled genomes (MAGs), which are further complicated by issues of contamination and high fragmentation. In contrast, long-read sequencing technologies, such as those from Oxford Nanopore Technologies (ONT) and PacBio HiFi, have significantly improved assembly quality. PacBio HiFi reads, known for their high accuracy, have enabled the resolution of numerous circular genomes from complex samples, while recent improvements in ONT technology have reduced error rates to approximately 1%, facilitating new methodological developments.

The introduction highlights several key assemblers developed for long-read metagenomic data, including MetaFlye, which was the first specifically designed for this purpose, and Hifiasm-meta, which excels in reconstructing complete, circularized genomes using HiFi reads. Furthermore, the introduction mentions the recent development of metaMDBG, which optimizes HiFi data assembly by significantly reducing memory usage and outperforming existing tools in generating near-complete MAGs and runtime efficiency. These advancements underscore the potential of long-read sequencing technologies to enhance our understanding of microbial diversity and functionality in various environments.

Methods

In the Methods section, the authors outline the procedural framework for the nanoMDBG approach, as illustrated in Figure 1. This framework encompasses a series of systematic steps designed to facilitate the implementation of nanoMDBG in their research. Each step is elaborated upon in subsequent sections, providing clarity on the methodologies employed and the rationale behind them. The detailed description aims to ensure reproducibility and to highlight the innovative aspects of the nanoMDBG technique.

Results

In this section, the authors present the results of their assembly evaluations, beginning with an analysis of assemblies derived solely from Oxford Nanopore Technologies (ONT) data. Following this, they compare the assemblies produced using both ONT and High-Fidelity (HiFi) data from identical samples. The section concludes with an assessment of assembly errors observed across both sequencing platforms, highlighting the differences in performance and accuracy between the two methodologies. This comprehensive evaluation aims to elucidate the strengths and weaknesses of each data type in genomic assembly processes.

Discussion

In this section, the authors present nanoMDBG, a novel metagenome assembly method tailored for Oxford Nanopore Technologies (ONT) long-read data, which enhances the existing metaMDBG assembler. The method employs a minimizer-space approach to correct reads and assemble metagenomes, significantly improving the reconstruction of metagenome-assembled genomes (MAGs), phages, and plasmids. Evaluation against state-of-the-art assemblers, including metaFlye and metaMDBG, reveals that nanoMDBG consistently outperforms these tools in terms of the number of near-complete MAGs, particularly in complex datasets. For instance, in the Human gut and Soil samples, nanoMDBG reconstructed substantially more near-complete MAGs and single-contig MAGs compared to its counterparts.

The authors also highlight that while ONT data shows promise, particularly in cost-effectiveness and scalability for metagenomic studies, challenges remain in achieving high coverage for complex communities like soil. The results indicate that only 37% of high-coverage genomes were converted into MAGs from the Soil dataset, compared to 85% from the Human gut, suggesting potential areas for further improvement. Overall, nanoMDBG represents a significant advancement in metagenomic assembly, enabling more comprehensive and cost-effective analyses of microbial communities, with implications for applications such as strain tracking and understanding evolutionary dynamics.

شارك: