تجميع جينوم على مستوى الكروموسوم للنبات الطبي الجبلي بيرغينيا بوربوراسنس (Saxifragaceae)
A Chromosome-level genome assembly of the alpine medicinal plant Bergenia purpurascens (Saxifragaceae)

شارك:
المجلة: Scientific Data، المجلد: 12، العدد: 1
DOI: https://doi.org/10.1038/s41597-025-04473-3
PMID: https://pubmed.ncbi.nlm.nih.gov/39837885
تاريخ النشر: 2025-01-21
المؤلف: Xue-Die Liu وآخرون
الموضوع الرئيسي: دراسات الكيمياء النباتية والنشاط الحيوي

نظرة عامة

تركز البحث على Bergenia purpurascens، وهو نوع متوطن جبلي من جبال الهيمالايا-هينغدوان، والذي يعتبر مهمًا لخصائصه الطبية واستخدامه الزخرفي. يقدم الدراسة تجميع جينومي عالي الجودة لـ B. purpurascens، يشمل 650.70 ميغابايت عبر 17 كروموسوم، مع طول سقالة N50 يبلغ 34.19 ميغابايت. كشفت توضيحات الجينوم عن 45,841 جينًا مشفرًا للبروتين و10,797 RNA غير مشفر، حيث تشكل التسلسلات المتكررة 56.72% من الجينوم. من المتوقع أن يعزز هذا الجينوم المرجعي فهم تخليق البرغين والتكيفات الجينية للنباتات الجبلية مع الظروف البيئية القاسية.

تواجه النباتات الجبلية، بما في ذلك B. purpurascens، تحديات شديدة مثل درجات الحرارة المنخفضة، والإشعاع فوق البنفسجي العالي، ومواسم النمو المحدودة. وقد سلطت الدراسات الجينومية السابقة على الأنواع الجبلية غير النموذجية الضوء على توسع عائلات الجينات المرتبطة بالتكيف الجبلي وتقلص جينات مقاومة الأمراض. ومع ذلك، فإن ندرة الجينومات المرجعية عالية الجودة قد أعاقت الرؤى الشاملة حول الآليات الجينية للتكيف في هذه الأنواع. نظرًا لبيئتها الفريدة وهشاشتها تجاه تغير المناخ، تعتبر B. purpurascens مرشحًا مثاليًا للتحقيق في الأسس الجينية للتكيف في النباتات الجبلية.

طرق

في هذه الدراسة، تم إجراء تسلسل الجينوم الكامل لنوع النبات B. purpurascens باستخدام عينات تم جمعها من جبل يولوونغ، مقاطعة يونان، الصين. تم استخراج الحمض النووي الجيني عالي الجودة وRNA الكلي من الأوراق باستخدام مجموعة Qiagen DNeasy Plant Mini Kit. تم بناء مكتبات القراءة القصيرة بحجم إدخال 350 نقطة أساس وتم تسلسلها على منصة Novaseq 6000، مما أسفر عن 50.71 غيغابايت من البيانات الخام، والتي غطت 79.91× من الجينوم. تم استخدام القراءات النظيفة (45.02 غيغابايت) لاستطلاعات الجينوم لتقييم حجم الجينوم، ومحتوى GC، والتغايرية.

بالإضافة إلى ذلك، تم إعداد مكتبات تسلسل PacBio باستخدام مجموعة تحضير القالب SMRTbell® express 2.0، مما أسفر عن 24.85 غيغابايت من البيانات النظيفة عالية الدقة (HiFi) مع قيمة N50 تبلغ 65.28 كيلوبايت ومتوسط طول قراءة يبلغ 20.31 كيلوبايت، تم إنتاجها على منصة PacBio Sequel II. تم تقطيع الحمض النووي الجيني إلى قطع بحجم 15 كيلوبايت قبل إعداد المكتبة. علاوة على ذلك، تم إعداد مكتبات Hi-C بأحجام إدخال تتراوح من 300 إلى 700 نقطة أساس وفقًا لمنهجية راو وآخرون وتم تسلسلها على منصة NovaSeq 6000، مما أنتج حوالي 72.85 غيغابايت من القراءات.

مناقشة

في هذا القسم، يوضح المؤلفون تجميع الجينوم على مستوى الكروموسوم وتوضيح *B. purpurascens*. استخدم التجميع الأولي للجينوم مكتبات إدخال قصيرة من Illumina وSOAP v.2.2115، مما أسفر عن تقدير حجم الجينوم بـ 634.52 ميغابايت، ومعدل تغايرية يبلغ 1.24%، ونسبة عناصر متكررة تبلغ 56.72%. تم تحقيق تجميع لاحق باستخدام قراءات طويلة عالية الدقة باستخدام Hifiasm v.0.16 وتدعيم عبر تقنيات Hi-C، مما حقق حجم تجميع نهائي قدره 650.70 ميغابايت، مع ربط 93.43% من الجينوم على 17 كروموسوم. أشار تحليل BUSCO إلى اكتمال عالٍ للتجميع، حيث تم تجميع 95.66% من 1,544 جين BUSCO بشكل كامل.

للتوضيح الجينومي، استخدم المؤلفون أدوات متنوعة، بما في ذلك MAKER v3.01.03 وRepeatMasker v4.1.2، مما أدى إلى التنبؤ بـ 45,841 جينًا مشفرًا للبروتين. تم تعزيز التوضيح الوظيفي من خلال المقارنات مع قواعد بيانات متعددة، مما أسفر عن نجاح توضيح 41,738 جينًا (91.05% من PCGs المتوقعة). بالإضافة إلى ذلك، حددت تنبؤات RNA غير المشفر 172 ميكرو RNA، و1,176 tRNA، و134 snRNA، و9,175 rRNA. تم إيداع بيانات تجميع الجينوم والتوضيح في عدة قواعد بيانات، مما يضمن الوصول إليها لمزيد من البحث. بشكل عام، تقدم الدراسة موردًا جينوميًا شاملاً لـ *B. purpurascens*، مع مقاييس تجميع وتوضيح عالية الجودة.

Journal: Scientific Data, Volume: 12, Issue: 1
DOI: https://doi.org/10.1038/s41597-025-04473-3
PMID: https://pubmed.ncbi.nlm.nih.gov/39837885
Publication Date: 2025-01-21
Author(s): Xue-Die Liu et al.
Primary Topic: Phytochemistry and Bioactivity Studies

Overview

The research focuses on Bergenia purpurascens, an alpine endemic species from the Himalaya-Hengduan Mountains (HHM), which is significant for its medicinal properties and ornamental use. The study presents a high-quality genome assembly of B. purpurascens, encompassing 650.70 Mb across 17 chromosomes, with an N50 scaffold length of 34.19 Mb. The genome annotation revealed 45,841 protein-coding genes and 10,797 noncoding RNAs, with repeat sequences constituting 56.72% of the genome. This reference genome is poised to enhance understanding of bergenin biosynthesis and the genetic adaptations of alpine plants to extreme environmental conditions.

Alpine plants, including B. purpurascens, face severe challenges such as low temperatures, high UV radiation, and limited growing seasons. Previous genomic studies on non-model alpine species have highlighted the expansion of gene families associated with alpine adaptation and the contraction of disease resistance genes. However, the scarcity of high-quality reference genomes has hindered comprehensive insights into the genetic mechanisms of adaptation in these species. Given its unique ecological niche and vulnerability to climate change, B. purpurascens serves as an ideal candidate for investigating the genetic underpinnings of adaptation in alpine flora.

Methods

In this study, whole-genome sequencing of the plant species B. purpurascens was conducted using samples collected from Yulong Mountain, Yunnan Province, China. High-quality genomic DNA and total RNA were extracted from the leaves using the Qiagen DNeasy Plant Mini Kit. Short-read libraries were constructed with an insert size of 350 bp and sequenced on the Novaseq 6000 platform, yielding 50.71 Gb of raw data, which covered 79.91× of the genome. The clean reads (45.02 Gb) were utilized for genome surveys assessing genome size, GC content, and heterozygosity.

Additionally, PacBio sequencing libraries were prepared using the SMRTbell® express template prep kit 2.0, resulting in 24.85 Gb of high-fidelity (HiFi) clean data with an N50 value of 65.28 kb and an average read length of 20.31 kb, generated on the PacBio Sequel II platform. Genomic DNA was sheared into 15 kb fragments prior to library preparation. Furthermore, Hi-C libraries with insert sizes ranging from 300 to 700 bp were prepared following the methodology of Rao et al. and sequenced on the NovaSeq 6000 platform, producing approximately 72.85 Gb of reads.

Discussion

In this section, the authors detail the chromosome-level genome assembly and annotation of *B. purpurascens*. Initial genome assembly utilized Illumina short-insert libraries and SOAP v.2.2115, resulting in an estimated genome size of 634.52 Mb, a heterozygosity rate of 1.24%, and a repeat element ratio of 56.72%. Subsequent assembly with HiFi long reads using Hifiasm v.0.16 and scaffolding via Hi-C techniques achieved a final assembly size of 650.70 Mb, with 93.43% of the genome anchored onto 17 chromosomes. The BUSCO analysis indicated a high completeness of the assembly, with 95.66% of the 1,544 BUSCO genes fully assembled.

For genome annotation, the authors employed various tools, including MAKER v3.01.03 and RepeatMasker v4.1.2, leading to the prediction of 45,841 protein-coding genes. Functional annotation was enhanced through comparisons with multiple databases, resulting in 41,738 genes (91.05% of predicted PCGs) being successfully annotated. Additionally, noncoding RNA predictions identified 172 microRNAs, 1,176 tRNAs, 134 snRNAs, and 9,175 rRNAs. The genome assembly and annotation data have been deposited in several databases, ensuring accessibility for further research. Overall, the study presents a comprehensive genomic resource for *B. purpurascens*, with high-quality assembly and annotation metrics.

شارك: