DOI: https://doi.org/10.1371/journal.pbio.3003659
PMID: https://pubmed.ncbi.nlm.nih.gov/41701781
تاريخ النشر: 2026-02-17
المؤلف: Yang Li وآخرون
الموضوع الرئيسي: آليات تخليق RNA والبروتين
نظرة عامة
تقدم البحث DRfold2، وهو إطار عمل جديد للتعلم العميق مصمم لتوقع بنية RNA من بيانات التسلسل. يدمج هذا الإطار نموذج لغة مركب RNA مدرب مسبقًا (RCLM) مع وحدة هيكلية لإزالة الضوضاء، مما يمكّن من التوقعات من البداية إلى النهاية. يتفوق DRfold2 بشكل كبير على الطرق الحالية الرائدة في كل من التوبولوجيا العالمية وتوقعات البنية الثانوية عبر اختبارات مرجعية متنوعة تشمل أنواعًا مختلفة. تُعزى التحسينات في الأداء إلى قدرة RCLM على التقاط أنماط التطور المشترك وفعالية عملية إزالة الضوضاء، مما يؤدي إلى زيادة تزيد عن 100% في دقة توقع الاتصال مقارنة بالأساليب السابقة.
بالإضافة إلى ذلك، يظهر DRfold2 توافقًا عاليًا مع AlphaFold3، مما يؤدي إلى تحسينات دقة ذات دلالة إحصائية عند دمجه في إطار تحسين. من خلال الجمع الفريد بين نمذجة اللغة المركبة، والتعلم القائم على إزالة الضوضاء، والتوجيه بالتعلم العميق بعد التحسين، يحدد DRfold2 مسارًا جديدًا لتقدم منهجيات توقع بنية RNA من البداية.
مقدمة
تسلط مقدمة هذه الورقة البحثية الضوء على الدور الحاسم لجزيئات RNA، وخاصة RNAs غير المشفرة، في وظائف خلوية متنوعة مثل تنظيم الجينات، والتحفيز، واستجابات الإجهاد. نظرًا لأهميتها، يتم استهداف RNAs بشكل متزايد لتطوير أدوية جزيئية صغيرة، خاصة في الأمراض التي تفتقر إلى أهداف بروتينية تقليدية. ومع ذلك، فإن التحدي المتمثل في تحديد هياكل RNA الذرية من التسلسلات الأولية قد زاد بسبب التراكم السريع لبيانات تسلسل RNA والافتقار إلى الهياكل المحلولة تجريبيًا. غالبًا ما تكون الطرق التجريبية التقليدية لتحديد بنية RNA مكلفة وغير عملية، مما يؤدي إلى الاعتماد المتزايد على الأساليب الحسابية لنمذجة الهياكل ثلاثية الأبعاد عالية الجودة.
تناقش الورقة تطور طرق توقع بنية RNA، من الأساليب المعتمدة على القوالب والأساليب من البداية إلى التقدمات الحديثة في التعلم العميق. بينما تواجه الطرق التقليدية قيودًا بسبب ندرة الهياكل المحلولة، أظهرت نماذج التعلم العميق مثل DeepFoldRNA وRhoFold وعدًا في توقع هياكل RNA. يقدم المؤلفون DRfold2، وهو إطار عمل جديد للتعلم العميق يستفيد من نموذج لغة مركب RNA مدرب مسبقًا (RCLM) لتعزيز تقريب الاحتمالية والتقاط إشارات التطور المشترك من تسلسلات RNA غير المراقبة. يدمج هذا الإطار تمثيلات تسلسلية وزوجية تتم معالجتها من خلال خط أنابيب قائم على Transformer، مصحوبًا بوحدة إزالة الضوضاء لتوقع البنية من البداية إلى النهاية. تشير نتائج المعايير إلى أن دمج ميزات التطور المشترك يحسن بشكل كبير من دقة توقعات بنية RNA من البداية، مما يمثل تقدمًا كبيرًا في هذا المجال.
طرق
تقيّم قسم الطرق في الورقة البحثية أداء DRfold2، وهو خط أنابيب جديد لتوقع بنية RNA، مقارنة بخمس طرق رائدة: RNAComposer وtrRosettaRNA وRhoFold وRoseTTAFoldNA وDeepFoldRNA، بالإضافة إلى AlphaFold3. تستخدم التقييمات درجة نمذجة القالب (TM-score) والانحراف المعياري الجذري (RMSD) لتقييم جودة الهياكل المتوقعة عبر عتبات هوية تسلسل متنوعة. يحقق DRfold2 باستمرار أعلى متوسط درجات TM، متفوقًا بشكل ملحوظ على DeepFoldRNA بنسبة 19.9% عند قطع هوية تسلسل 80%، مع قيمة p ذات دلالة إحصائية تبلغ 6.37E-04. بالإضافة إلى ذلك، يظهر DRfold2 قيم RMSD متفوقة مقارنة بجميع طرق التحكم، مما يدل على قوته في توقعات البنية الثلاثية العالمية.
كما يبرز القسم نقاط القوة التكميلية لـ DRfold2 وAlphaFold3، خاصة في التعامل مع هياكل RNA ذات العقد الزائفة، حيث يظهر DRfold2 درجات TM أعلى ولكن معدلات استرداد أقل للعقد الزائفة مقارنة بـ AlphaFold3. يعزز دمج تكوينات AlphaFold3 في وحدة تحسين DRfold2 جودة التوقع، مما يؤدي إلى متوسط درجة TM تبلغ 0.368 وRMSD قدره 14.02 Å، وهو أفضل بكثير من AlphaFold3 بمفرده. على الرغم من درجات TM المتواضعة بشكل عام، يتوقع DRfold2 بشكل موثوق توبولوجيات هيكلية ذات مغزى عبر تسلسلات RNA متنوعة، متفوقًا على RNAComposer ويظهر معدل فشل منخفض في التقاط الطيات الهيكلية الصحيحة. تسهم التحسينات في بنية DRfold2، بما في ذلك نموذج لغة RNA المركب ووحدة هيكل إزالة الضوضاء، في تحسين أدائه مقارنة بسابقه، DRfold.
نتائج
يستعرض قسم النتائج خط أنابيب DRfold2 لتوقع بنية RNA، والذي يبدأ بإدماج تسلسل RNA في تمثيلات تسلسلية وزوجية باستخدام نموذج لغة سياقية متكررة مدرب مسبقًا (RCLM). يتفوق هذا النموذج في التعرف على أنماط التسلسل بفضل تدريبه على بيانات تسلسل غير مراقبة واسعة النطاق عبر نهج تعظيم الاحتمالية المركبة. تتم معالجة التمثيلات الناتجة من خلال كتل RNA Transformer، التي تسهل التحولات اللازمة لطي بنية RNA. بعد ذلك، تنتج وحدة هيكل RNA لإزالة الضوضاء (DRSM) تكوينات RNA بطريقة من البداية إلى النهاية، ويتم استخدام بروتوكول اختيار وتكرير التكوين (CSOR) لتحسين النماذج النهائية لـ RNA من مجموعة من النماذج المولدة.
لتقييم أداء DRfold2، تم إنشاء مجموعة بيانات اختبار مستقلة، تتكون من 41 هيكل RNA بأطوال أقل من 400 نيوكليوتيد. تتضمن هذه المجموعة أهداف RNA-Puzzles، وأهداف RNA من CASP15، وهياكل RNA فقط التي تم إصدارها مؤخرًا من بنك بيانات البروتين، مع استبعاد RNAs الاصطناعية الكبيرة من CASP15 بسبب طبيعتها غير العادية. تتكون مجموعة التدريب من هياكل RNA التي تم إصدارها قبل عام 2024، مما يضمن عدم تضمين أي RNAs تشارك أكثر من 80% من هوية التسلسل مع RNAs الاختبار. كلا المجموعتين متاحتان للجمهور لمزيد من البحث والتحقق.
مناقشة
تسلط قسم المناقشة في الورقة البحثية الضوء على قدرات وأداء DRfold2 في توقع بنية RNA، خاصة للأهداف ذات الطيات الجديدة ومستويات مختلفة من المعلومات التطورية. يظهر DRfold2 أداءً قويًا، حيث يحقق متوسط درجة TM تبلغ 0.309 على 25 هدفًا ذو طيات جديدة، والتي، على الرغم من كونها أقل من أدائه على الطيات المعروفة (0.350)، تشير إلى قدرته على التعميم بما يتجاوز حفظ الهياكل التي تم رؤيتها سابقًا. من الجدير بالذكر أن DRfold2 يتفوق على AlphaFold3 في هذه المجموعة الفرعية، حيث حصل على درجة TM تبلغ 0.309 مقارنة بـ 0.281 لـ AlphaFold3، وانحراف معياري أقل بكثير، مما يظهر فعاليته في استخراج المبادئ الهيكلية من التسلسلات.
بالإضافة إلى ذلك، تناقش الورقة تأثير عمق المحاذاة المتعددة للتسلسل (MSA) على دقة التوقع. يظهر DRfold2 أداءً متفوقًا في السيناريوهات التي تحتوي على عدد قليل أو لا يوجد فيها نظائر، حيث يحقق درجة TM تبلغ 0.300 للأهداف التي لا تحتوي على نظائر قابلة للاكتشاف، مقارنة بـ 0.287 لـ AlphaFold3. يشير هذا إلى أن تدريب DRfold2 على ميزات التسلسل الفردي المثرى بإشارات التطور المشترك يسمح له بإجراء توقعات موثوقة حتى في غياب معلومات تطورية واسعة. علاوة على ذلك، يستخدم DRfold2 استراتيجية تحسين هجينة تعزز الواقعية الفيزيائية للهياكل المتوقعة، محققًا تحسينات كبيرة في مقاييس الجودة الهندسية. بشكل عام، تؤكد النتائج على نقاط قوة DRfold2 في توقع بنية RNA، خاصة في السيناريوهات التي تفتقر إلى تسلسلات نظيرية واسعة، وإمكاناته للتكامل مع نماذج أخرى مثل AlphaFold3 من أجل دقة محسنة.
DOI: https://doi.org/10.1371/journal.pbio.3003659
PMID: https://pubmed.ncbi.nlm.nih.gov/41701781
Publication Date: 2026-02-17
Author(s): Yang Li et al.
Primary Topic: RNA and protein synthesis mechanisms
Overview
The research presents DRfold2, a novel deep learning framework designed for RNA structure prediction from sequence data. This framework integrates a pre-trained RNA Composite Language Model (RCLM) with a denoising structure module, enabling end-to-end predictions. DRfold2 significantly outperforms existing state-of-the-art methods in both global topology and secondary structure predictions across various benchmark tests involving diverse species. The enhancements in performance are attributed to the RCLM’s capacity to capture coevolutionary patterns and the effectiveness of the denoising process, resulting in over a 100% increase in contact prediction precision compared to prior approaches.
Additionally, DRfold2 shows high compatibility with AlphaFold3, yielding statistically significant accuracy improvements when incorporated into an optimization framework. By uniquely combining composite language modeling, denoising-based learning, and deep learning-guided post-optimization, DRfold2 sets a new trajectory for advancing ab initio RNA structure prediction methodologies.
Introduction
The introduction of this research paper highlights the critical role of RNA molecules, particularly noncoding RNAs, in various cellular functions such as gene regulation, catalysis, and stress responses. Given their significance, RNAs are increasingly targeted for small-molecule drug development, especially in diseases lacking traditional protein targets. However, the challenge of determining atomistic RNA structures from primary sequences has intensified due to the rapid accumulation of RNA sequence data and the limited availability of experimentally resolved structures. Traditional experimental methods for RNA structure determination are often costly and impractical, leading to a growing reliance on computational approaches for high-quality 3D structure modeling.
The paper discusses the evolution of RNA structure prediction methods, from template-based and ab initio approaches to recent advancements in deep learning. While traditional methods face limitations due to the scarcity of solved structures, deep learning models like DeepFoldRNA and RhoFold have shown promise in predicting RNA structures. The authors introduce DRfold2, a novel deep learning framework that leverages a pretrained RNA Composite Language Model (RCLM) to enhance likelihood approximation and capture co-evolutionary signals from unsupervised RNA sequences. This framework integrates sequential and pairwise representations processed through a Transformer-based pipeline, coupled with a denoising module for end-to-end structure prediction. Benchmark results indicate that the incorporation of co-evolution features significantly improves the accuracy of ab initio RNA structure predictions, marking a substantial advancement in the field.
Methods
The methods section of the research paper evaluates the performance of DRfold2, a novel RNA tertiary structure prediction pipeline, against five leading methods: RNAComposer, trRosettaRNA, RhoFold, RoseTTAFoldNA, and DeepFoldRNA, as well as AlphaFold3. The evaluation employs Template Modeling Score (TM-score) and Root Mean Square Deviation (RMSD) to assess the quality of predicted structures across various sequence identity thresholds. DRfold2 consistently achieves the highest average TM-scores, notably outperforming DeepFoldRNA by 19.9% at an 80% sequence identity cutoff, with a statistically significant p-value of 6.37E-04. Additionally, DRfold2 demonstrates superior RMSD values compared to all control methods, indicating its robustness in global tertiary structure predictions.
The section also highlights the complementary strengths of DRfold2 and AlphaFold3, particularly in handling RNA structures with pseudoknots, where DRfold2 shows higher average TM-scores but lower recovery rates for pseudoknots compared to AlphaFold3. The integration of AlphaFold3 conformations into DRfold2’s optimization module enhances prediction quality, yielding an average TM-score of 0.368 and RMSD of 14.02 Å, significantly better than AlphaFold3 alone. Despite the overall modest TM-scores, DRfold2 reliably predicts meaningful structural topologies across diverse RNA sequences, outperforming RNAComposer and demonstrating a low failure rate in capturing correct structural folds. The advancements in DRfold2’s architecture, including the RNA Composite Language Model and Denoising Structure Module, contribute to its improved performance over its predecessor, DRfold.
Results
The results section outlines the DRfold2 pipeline for RNA structure prediction, which begins with embedding an RNA sequence into sequence and pair representations using a pretrained Recurrent Contextual Language Model (RCLM). This model excels in recognizing sequence patterns due to its training on extensive unsupervised sequence data via a composite likelihood maximization approach. The representations generated are then processed through RNA Transformer Blocks, which facilitate the necessary transformations for RNA structure folding. Subsequently, the Denoising RNA Structure Module (DRSM) generates RNA conformations in an end-to-end manner, and a post-processing Conformation Selection and Optimization Refinement (CSOR) protocol is employed to refine the final RNA models from a pool of generated decoys.
To evaluate the performance of DRfold2, an independent test dataset was created, comprising 41 RNA structures with lengths under 400 nucleotides. This dataset includes RNA-Puzzles targets, CASP15 RNA targets, and recently released RNA-only structures from the Protein Data Bank, while excluding large synthetic RNAs from CASP15 due to their atypical nature. The training set consists of RNA structures released prior to 2024, ensuring that no RNAs sharing over 80% sequence identity with the test RNAs are included. Both datasets are publicly accessible for further research and validation.
Discussion
The discussion section of the research paper highlights the capabilities and performance of DRfold2 in RNA structure prediction, particularly for targets with novel folds and varying levels of evolutionary information. DRfold2 demonstrates robust performance, achieving an average TM-score of 0.309 on 25 novel-fold targets, which, while lower than its performance on known folds (0.350), indicates its ability to generalize beyond memorization of previously seen structures. Notably, DRfold2 outperforms AlphaFold3 on this subset, with a TM-score of 0.309 compared to AlphaFold3’s 0.281, and a significantly lower RMSD, showcasing its effectiveness in extracting structural principles from sequences.
Additionally, the paper discusses the impact of multiple sequence alignment (MSA) depth on prediction accuracy. DRfold2 shows superior performance in scenarios with few or no homologs, achieving a TM-score of 0.300 for targets without detectable homologs, compared to AlphaFold3’s 0.287. This suggests that DRfold2’s training on single-sequence features enriched with co-evolutionary signals allows it to make reliable predictions even in the absence of extensive evolutionary information. Furthermore, DRfold2 employs a hybrid refinement strategy that enhances the physical realism of predicted structures, achieving significant improvements in geometric quality metrics. Overall, the findings underscore DRfold2’s strengths in RNA structure prediction, particularly in scenarios lacking extensive homologous sequences, and its potential for integration with other models like AlphaFold3 for enhanced accuracy.
