بناء رسومات الجينوم الشامل
Building pangenome graphs

شارك:
المجلة: Nature Methods، المجلد: 21، العدد: 11
DOI: https://doi.org/10.1038/s41592-024-02430-3
PMID: https://pubmed.ncbi.nlm.nih.gov/39433878
تاريخ النشر: 2024-10-21
المؤلف: Erik Garrison وآخرون
الموضوع الرئيسي: المعلوماتية الحيوية والشبكات الجينومية

نظرة عامة

يقدم هذا القسم نظرة عامة على مُنشئ رسم الجينوم الشامل (PGGB)، وهو خط أنابيب جديد مصمم لبناء رسوم الجينوم الشامل التي تمثل بشكل شامل التنوع الجيني بين عدة جينومات مرجعية. على عكس الطرق الحالية التي تستبعد إما التسلسلات المعقدة أو تعتمد على جينوم مرجعي واحد، يستخدم PGGB استراتيجية محاذاة شاملة. تتيح هذه الطريقة تحديد التنوع الجيني، وقياس مستويات الحفظ، واكتشاف أحداث إعادة التركيب، واستنتاج العلاقات النشوء والتطور داخل البيانات الجينومية.

من خلال معالجة قيود تقنيات بناء رسوم الجينوم الشامل الحالية، يوفر PGGB إطارًا أكثر شمولية وحيادية لتحليل التنوع الجينومي. هذه التقدمات ضرورية لتعزيز فهمنا للعمليات التطورية والهندسة الجينية للسكان.

طرق

في هذا القسم، يوضح المؤلفون المنهجية الأساسية لخوارزمية SMOOTHXG، التي تعد جزءًا لا يتجزأ من إطار PGGB لتطبيع الرسوم وتبسيطها. تعمل SMOOTHXG من خلال عدة تمريرات على الجينوم الشامل، مما يعيد تشكيل الرسم بشكل فعال لتخفيف مشكلات التعقيد المحلي وعدم المحاذاة التي كانت شائعة في طرق بناء الرسوم المعتمدة على المحاذاة الشاملة السابقة. غالبًا ما أدت هذه الطرق السابقة إلى أنماط رسومية معقدة ومتكررة على مقاييس أصغر، بالإضافة إلى تحديات تتعلق بالازدواجية في الرسوم وانخفاض حساسية المحاذاة بسبب تصفية المطابقات.

علاوة على ذلك، يحدد المؤلفون منهجية التقييم الخاصة بهم لتجارب التحقق المتبادل، حيث يتم مقارنة الرسوم الناتجة عن PGGB مع تعدد أشكال النوكليوتيدات المفردة (SNPs) التي تم تحديدها باستخدام MUMMER4. تعمل هذه التحليل المقارن على التحقق من فعالية خوارزمية SMOOTHXG في تحسين جودة الرسم ودقة المحاذاة.

مناقشة

في هذا القسم، يوضح المؤلفون المنهجيات المستخدمة لبناء رسوم الجينوم الشامل عبر أنواع مختلفة، بما في ذلك *Arabidopsis thaliana*، *Escherichia coli*، *Homo sapiens*، *Mus musculus*، وغيرها. يصفون معايير اختيار التجميعات الجينومية، مؤكدين على أهمية حل الجينومات على مستوى الكروموسوم وضمان استبعاد القطع غير الكروموسومية. كما يقدم المؤلفون خوارزمية SMOOTHXG، التي تعالج الرسوم الخام لـ SEQWISH لتعزيز قابليتها للاستخدام من خلال فرز وتطبيع هيكل الرسم، مما يحسن تمثيل التنوع الجينومي. يتضمن ذلك سلسلة من خطوات المعالجة المسبقة، بما في ذلك خوارزمية الانحدار العشوائي الموجه بالمسار (PG-SGD) لترتيب العقد وتحليل إحصائي لاكتشاف التكرارات المتغيرة العدد (VNTRs).

علاوة على ذلك، يناقش المؤلفون التحقق من طرق بناء رسوم الجينوم الشامل واستدعاء المتغيرات من خلال التحقق المتبادل ضد أداة NUCMER. يحددون نهجًا منهجيًا لاستخراج التسلسلات، وتحديد المتغيرات، وتقييم أداء طريقتهم المعتمدة على الرسم مقارنة بتقنيات المحاذاة التقليدية. تشير النتائج إلى أنه بينما يلتقط PGGB المتغيرات الصغيرة بشكل فعال، لا تزال هناك تحديات، خاصة في المناطق ذات الكثافات العالية من المتغيرات النوكليوتيدية المفردة المتقاربة (SNVs). كما يبرز المؤلفون الحاجة إلى تحسين مستمر في إجراءات القياس لاستدعاء المتغيرات الهيكلية، مؤكدين على تعقيد تمثيل العلاقات الجينومية بدقة في رسوم الجينوم الشامل.

Journal: Nature Methods, Volume: 21, Issue: 11
DOI: https://doi.org/10.1038/s41592-024-02430-3
PMID: https://pubmed.ncbi.nlm.nih.gov/39433878
Publication Date: 2024-10-21
Author(s): Erik Garrison et al.
Primary Topic: Bioinformatics and Genomic Networks

Overview

The section presents an overview of the PanGenome Graph Builder (PGGB), a novel pipeline designed to construct pangenome graphs that comprehensively represent genetic variation among multiple reference genomes. Unlike existing methods that either exclude complex sequences or rely on a single reference genome, PGGB employs an all-to-all alignment strategy. This approach enables the identification of genetic variation, measurement of conservation levels, detection of recombination events, and inference of phylogenetic relationships within the genomic data.

By addressing the limitations of current pangenome graph construction techniques, PGGB offers a more inclusive and unbiased framework for analyzing genomic diversity. This advancement is crucial for enhancing our understanding of evolutionary processes and the genetic architecture of populations.

Methods

In this section, the authors detail the methodology underlying the SMOOTHXG algorithm, which is integral to the PGGB framework for graph normalization and simplification. SMOOTHXG operates through multiple passes over the pangenome, effectively reshaping the graph to mitigate issues of local complexity and underalignment that were prevalent in prior all-vs-all alignment-based graph construction methods. These earlier approaches often led to intricate, looping graph motifs at smaller scales, as well as challenges related to graph redundancy and diminished alignment sensitivity due to match filtering.

Furthermore, the authors outline their evaluation methodology for cross-validation experiments, wherein the graphs generated by PGGB are compared against single nucleotide polymorphisms (SNPs) identified using MUMMER4. This comparative analysis serves to validate the effectiveness of the SMOOTHXG algorithm in enhancing graph quality and alignment accuracy.

Discussion

In this section, the authors detail the methodologies employed for constructing pangenome graphs across various species, including *Arabidopsis thaliana*, *Escherichia coli*, *Homo sapiens*, *Mus musculus*, and others. They describe the selection criteria for genomic assemblies, emphasizing the importance of resolving genomes at the chromosome level and ensuring the exclusion of non-chromosomal contigs. The authors also introduce the SMOOTHXG algorithm, which processes raw SEQWISH graphs to enhance their usability by sorting and normalizing the graph structure, thereby improving the representation of genomic variation. This involves a series of preprocessing steps, including a Path-Guided Stochastic Gradient Descent (PG-SGD) algorithm for node ordering and a statistical analysis for detecting Variable Number Tandem Repeats (VNTRs).

Furthermore, the authors discuss the validation of their pangenome graph construction and variant calling methods through cross-validation against the NUCMER tool. They outline a systematic approach for extracting sequences, identifying variants, and evaluating the performance of their graph-based method compared to traditional alignment techniques. The results indicate that while PGGB effectively captures small variants, challenges remain, particularly in regions with high densities of closely spaced single nucleotide variants (SNVs). The authors also highlight the need for ongoing refinement in benchmarking procedures for structural variant calling, emphasizing the complexity of accurately representing genomic relationships in pangenome graphs.

شارك: