FLOWR: مطابقة التدفق لتوليد الجزيئات الجديدة القائمة على التفاعل والشظايا مع مراعاة الهيكل
FLOWR: flow matching for structure-aware de novo, interaction- and fragment-based ligand generation

شارك:
المجلة: Nature Computational Science، المجلد: 6، العدد: 6
DOI: https://doi.org/10.1038/s43588-026-00998-8
PMID: https://pubmed.ncbi.nlm.nih.gov/42209794
تاريخ النشر: 2026-05-28
المؤلف: Julian Cremer وآخرون
الموضوع الرئيسي: طرق اكتشاف الأدوية الحاسوبية

طرق

في هذا القسم، يحدد المؤلفون الإطار المنهجي الذي يدعم نهج مطابقة التدفق المستخدم في FLOWR. تستند المنهجية إلى التدفقات الطبيعية المستمرة (CNFs) التي يتم تدريبها من خلال مطابقة التدفق الشرطية (CFM)، مما يدمج أهداف الانحدار المستقرة المميزة لنماذج الانتشار مع عملية استدلال حتمية فعالة خالية من المحاكاة. يقوم المؤلفون بتنفيذ مطابقة التدفق للنقل الأمثل (OT-CFM) لتبسيط وتثبيت بناء التدفق من خلال تقليل تكاليف النقل بين التوزيعات المصدر والهدف.

علاوة على ذلك، يعزز دمج مطابقة التدفق المتساوي النموذج من خلال الاستفادة من التماثلات الدورانية والترجمية الموجودة في الأنظمة الجزيئية. يؤدي ذلك إلى تدفقات لا تتمتع فقط بمسارات تكامل أقصر ولكن أيضًا تظهر كفاءة محسنة في العينة ودمج طبيعي للتماثلات الفيزيائية. هذه الميزات ضرورية للتوليد الدقيق لمجمعات البروتين-الليغاند الجيومترية الصالحة. يختتم القسم بوصف تفصيلي لكل مكون وتكاملها ضمن إطار النموذج العام.

نتائج

في هذا القسم، يقدم المؤلفون FLOWR، وهو نموذج توليدي قائم على التدفق مصمم للتوليد من الصفر لليغاندات المشروطة على أشكال جيوب البروتين المحددة وميزات الجيوب-الليغاند المرغوبة. يستفيد النموذج من مجموعة بيانات تتكون من أزواج من الليغاندات، جيوب البروتين، ومصفوفات ميزات ذرية اختيارية. يبني FLOWR على بنية SEMLA، ويشمل مشفر الجيب ووحدة انتباه متساوي لتسهيل التوليد الشرطي، مما يعزز قدرة النموذج على التقاط الفروق الهندسية والكيميائية المعقدة لتفاعلات البروتين-الليغاند.

كما يقدم المؤلفون FLOWR.MULTI، وهو امتداد لـ FLOWR يسمح بالتوليد الشرطي متعدد الأغراض، مما يمكّن من تصميم ليغاندات تتوافق مع ملفات التفاعل المحددة مسبقًا وهياكل كيميائية معينة دون الحاجة إلى إعادة تدريب النموذج. على الرغم من التقدم الذي يقدمه FLOWR، يعترف المؤلفون بالتحديات في تقييم منهجيات تصميم الأدوية المعتمدة على الهيكل (SBDD)، خاصة بسبب مشكلات جودة البيانات في مجموعات البيانات المرجعية مثل CROSSDOCKED2020، والتي قد تؤدي إلى تدريب النماذج على سيناريوهات ارتباط غير مثالية. يستخدم نموذج FLOWR مخطط مطابقة تدفق مختلط مستمر وفئوي لتوليد الليغاندات من خلال حل معادلة تفاضلية عادية (ODE) لتدفق الاحتمالات، مما يضمن توليد ليغاندات بكفاءة ودقة مع معالجة قيود الأساليب القائمة على الانتشار الموجودة.

نقاش

في هذا القسم، يناقش المؤلفون التقدم المحرز في نموذج FLOWR لتوليد الليغاندات، مع التركيز على كفاءته ودقته في توليد الميزات الجزيئية. يستخدم النموذج وحدة تغذية أمامية متساوية البوابات ويشمل تضمينات الروابط في طبقات الانتباه الذاتي، مما يعزز كل من الصلاحية والكفاءة الحسابية. يعتمد FLOWR نهج نمذجة مشتركة للميزات الجزيئية المستمرة والقطعية، مما يسمح بالتنبؤ المباشر بالشحنات الرسمية لليغاندات والتنقيح التكراري لعينات الليغاندات الضوضائية من خلال حقول متجهة متعلمة. يوسع إدخال FLOWR.MULTI قدرات النموذج لدعم مهام التوليد المختلفة، مثل القفز على السقالات والعينات القائمة على القطع، والتي تعتبر حاسمة لعمليات اكتشاف الأدوية.

كما يقدم المؤلفون SPINDR، مجموعة بيانات مصقولة تهدف إلى تحسين جودة بيانات تفاعل البروتين-الليغاند، ومعالجة المشكلات الموجودة في مجموعات البيانات الحالية مثل PDBBind. تتكون SPINDR من 35,666 مجمع بروتين-ليغاند عالي الجودة، مما يوفر معيارًا قويًا لتقييم النماذج التوليدية. يتم مقارنة أداء FLOWR مع أساليب توليد أخرى، مما يظهر نتائج متفوقة من حيث مقاييس الصلاحية، طاقة الإجهاد، وسرعة الاستدلال. من الجدير بالذكر أن FLOWR يحقق تحسينًا كبيرًا في جودة توليد الليغاندات، متفوقًا على نموذج PILOT عبر مقاييس متعددة، بما في ذلك درجات AutoDock-Vina والأمان الهيكلي. تشير النتائج إلى أن FLOWR لا يعزز فقط توليد الليغاندات ولكنه يقدم أيضًا موردًا قيمًا للبحث المستقبلي في تصميم الأدوية المعتمد على الهيكل.

Journal: Nature Computational Science, Volume: 6, Issue: 6
DOI: https://doi.org/10.1038/s43588-026-00998-8
PMID: https://pubmed.ncbi.nlm.nih.gov/42209794
Publication Date: 2026-05-28
Author(s): Julian Cremer et al.
Primary Topic: Computational Drug Discovery Methods

Methods

In this section, the authors outline the methodological framework underpinning the flow matching approach utilized in FLOWR. The methodology is based on continuous normalizing flows (CNFs) that are trained through conditional flow matching (CFM), integrating the stable regression objectives characteristic of diffusion models with a simulation-free, efficient deterministic inference process. The authors specifically implement optimal transport CFM (OT-CFM) to simplify and stabilize the flow construction by minimizing transport costs between source and target distributions.

Furthermore, the integration of equivariant flow matching enhances the model by leveraging the rotational and translational symmetries present in molecular systems. This results in flows that not only have shorter integration paths but also exhibit improved sampling efficiency and a natural incorporation of physical symmetries. These features are crucial for the accurate generation of geometrically valid protein-ligand complexes. The section concludes with a detailed description of each component and their integration within the overall model framework.

Results

In this section, the authors introduce FLOWR, a flow-based generative model designed for the de novo generation of ligands conditioned on specific protein pocket geometries and desired pocket-ligand features. The model leverages a dataset comprising tuples of ligands, protein pockets, and optional atomic feature matrices. FLOWR builds upon the SEMLA architecture, incorporating a pocket encoder and an equivariant cross-attention module to facilitate conditional generation, thereby enhancing the model’s ability to capture the intricate geometric and chemical nuances of protein-ligand interactions.

The authors also present FLOWR.MULTI, an extension of FLOWR that allows for multi-purpose conditional generation, enabling the design of ligands that adhere to predefined interaction profiles and specific chemical substructures without the need for model retraining. Despite the advancements offered by FLOWR, the authors acknowledge challenges in evaluating structure-based drug design (SBDD) methodologies, particularly due to data quality issues in benchmark datasets like CROSSDOCKED2020, which may lead to models being trained on non-ideal binding scenarios. The FLOWR model employs a mixed continuous and categorical flow matching scheme to generate ligands by solving a probability-flow ordinary differential equation (ODE), ensuring efficient and accurate ligand generation while addressing the limitations of existing diffusion-based approaches.

Discussion

In this section, the authors discuss the advancements made in the FLOWR model for ligand generation, emphasizing its efficiency and accuracy in generating molecular features. The model utilizes a gated equivariant feed-forward module and incorporates bond embeddings into self-attention layers, which enhances both validity and computational efficiency. FLOWR employs a joint modeling approach for continuous and discrete molecular features, allowing for the direct prediction of ligand formal charges and the iterative refinement of noisy ligand samples through learned vector fields. The introduction of FLOWR.MULTI extends the model’s capabilities to support various generation tasks, such as scaffold hopping and fragment-based sampling, which are crucial for drug discovery processes.

The authors also present SPINDR, a refined dataset aimed at improving the quality of protein-ligand interaction data, addressing issues found in existing datasets like PDBBind. SPINDR consists of 35,666 high-quality protein-ligand complexes, providing a robust benchmark for evaluating generative models. The performance of FLOWR is compared against other generative methods, demonstrating superior results in terms of validity metrics, strain energy, and inference speed. Notably, FLOWR achieves a significant improvement in ligand generation quality, outperforming the PILOT model across multiple metrics, including AutoDock-Vina scores and structural fidelity. The findings suggest that FLOWR not only enhances ligand generation but also offers a valuable resource for future research in structure-based drug design.

شارك: