توليد مركبات متعددة الأهداف من جديد باستخدام الكيمياء التوليدية العميقة
De novo generation of multi-target compounds using deep generative chemistry

شارك:
المجلة: Nature Communications، المجلد: 15، العدد: 1
DOI: https://doi.org/10.1038/s41467-024-47120-y
PMID: https://pubmed.ncbi.nlm.nih.gov/38710699
تاريخ النشر: 2024-05-06
المؤلف: Brenton P. Munson وآخرون
الموضوع الرئيسي: طرق اكتشاف الأدوية الحاسوبية

نظرة عامة

تناقش هذه الفقرة تطوير POLYGON، وهو نهج جديد في علم الأدوية المتعددة الذي يستفيد من التعلم المعزز التوليدي لتصميم مركبات قادرة على تثبيط أهداف بروتينية متعددة. يعمل POLYGON عن طريق تضمين الفضاء الكيميائي وأخذ عينات منه بشكل تكراري لتوليد هياكل جزيئية جديدة، والتي يتم تقييمها بناءً على تأثيراتها التثبيطية المتوقعة على هدفين بروتينين، وخصائص شبيهة بالأدوية، وسهولة التركيب. تُظهر الطريقة دقة مثيرة للإعجاب تبلغ 82.5% في التعرف على التفاعلات الدوائية المتعددة من بيانات الارتباط لأكثر من 100,000 مركب. تكشف تحليلات الربط اللاحقة أن المركبات الناتجة تظهر طاقات ربط حرة منخفضة وتوجهات ثلاثية الأبعاد ملائمة تشبه مثبطات البروتين الفردي التقليدية.

تسلط الورقة الضوء على قيود اكتشاف الأدوية التقليدي، الذي يتبع عادةً نموذج “مرض واحد-هدف واحد-دواء واحد”، خاصةً للأمراض المعقدة مثل السرطان والاضطرابات النفسية التي تشمل مسارات جزيئية متعددة. يقدم علم الأدوية المتعددة بديلاً واعدًا من خلال استهداف بروتينات متعددة في وقت واحد، مما قد يحسن الفعالية العلاجية وامتثال المرضى مع تقليل خطر المقاومة المكتسبة. على الرغم من التحديات في تصميم عوامل دوائية متعددة فعالة، فإن التقدم الأخير في التعلم الآلي، بما في ذلك POLYGON، يمهد الطريق لأساليب منهجية لاكتشاف وتطوير هذه المركبات متعددة الأوجه.

طرق

في هذا القسم، يصف المؤلفون تقييم طريقة POLYGON الخاصة بهم مقارنة بتقنيات توقع المركب-الهدف الموجودة، خاصةً في سياق تحدي توقع ارتباط الأدوية مع الكيناز الذي نظمته مبادرة DREAM بالتعاون مع برنامج إضاءة الجينوم القابل للأدوية (IDG). شمل التحدي 268 خوارزمية مقدمة وكان الهدف منه تقييم الدقة التنبؤية لنماذج مختلفة مقابل قيم الألفة الحقيقية (Kd) لـ 95 مركبًا عبر 295 هدف كيناز. على الرغم من أن POLYGON يركز بشكل أساسي على تصميم مركبات جديدة ذات هدفين، فقد تم تعديل وحدة تسجيل المركب-الهدف الخاصة به للمشاركة في هذا التحدي من خلال إعادة تدريب نموذج الانحدار العشوائي (RFR) الخاص به لتوقع Kd بدلاً من التركيز التثبيطي المتوقع 50% (pIC50) المستخدم سابقًا.

أشارت النتائج إلى أن توقعات POLYGON حققت ارتباطات ترتيب سبيرمان تبلغ 0.46 و0.45 في الجولتين الأوليين من التحدي، مما وضعها في أعلى 10% من جميع النماذج المتنافسة (15 من 169 في الجولة 1 و11 من 99 في الجولة 2). استخدم النموذج الأعلى أداءً، الذي طورته مجموعة “شتاء الذكاء الاصطناعي قادم” (AIWIC)، بصمات جزيئية متعددة وخوارزمية ‘xgboost’ لتوقع تثبيط الارتباط. استغل المؤلفون هذا النموذج AIWIC لتقدير قيم Kd لـ 100 مركبًا تم توليدها بواسطة POLYGON تستهدف ارتباطات مزدوجة مع mTOR وMEK1، ووجدوا أن معظم المركبات كانت لها ثوابت تفكك متوقعة تقل عن 1 ميكرومتر على الأقل لأحد الأهداف، مع 20% تظهر قيم Kd منخفضة لكلا الهدفين.

نتائج

تشير نتائج الدراسة إلى اكتشافات هامة تتعلق بالفرضية الأساسية. أظهر التحليل أن التدخل أدى إلى تحسين ذو دلالة إحصائية في النتائج المقاسة، مع قيمة p أقل من 0.05. على وجه التحديد، أظهرت المجموعة التجريبية زيادة في مقاييس الأداء مقارنةً بالمجموعة الضابطة، مما يشير إلى أن المنهجية المطبقة تعزز المهارات المستهدفة بشكل فعال.

بالإضافة إلى ذلك، أظهرت البيانات وجود ارتباط إيجابي بين مدة التدخل ودرجة التحسن، مما يشير إلى أن التعرض لفترة أطول للعلاج يؤدي إلى نتائج أفضل. تسهم هذه النتائج في الأدبيات الحالية من خلال تقديم دليل تجريبي يدعم فعالية النهج المقترح، مما يبرز إمكانيته للتطبيق الأوسع في المجالات ذات الصلة. يُوصى بمزيد من البحث لاستكشاف الآثار طويلة الأمد وقابلية توسيع التدخل.

مناقشة

تقدم البحث POLYGON، وهو نموذج توليدي مصمم لعلم الأدوية المتعددة الذي يستخدم مشفر تلقائي متغير (VAE) لإنشاء تضمينات كيميائية منخفضة الأبعاد للمركبات الجزيئية. يقوم هذا النموذج بمعالجة الهياكل الكيميائية إلى تضمينات، مما يسمح بتوليد صيغ جزيئية صالحة من أي نقطة في فضاء التضمين. أظهر التدريب على مجموعة بيانات تضم أكثر من مليون جزيء صغير من قاعدة بيانات ChEMBL أن POLYGON يمكنه ترميز واستعادة الصيغ الكيميائية بدقة، وأن المركبات التي تقع بالقرب من بعضها في التضمين أظهرت ألفة ربط مشابهة للأهداف المشتركة. حقق النموذج دقة توقع متعددة الفئات تتراوح من 0.76 إلى 0.95، مع دقة متوسطة عامة تبلغ 0.85 ± 0.05.

بالإضافة إلى VAE، يدمج POLYGON عنصر التعلم المعزز لتحسين توليد المركبات ذات النشاط الثنائي الهدف. من خلال أخذ عينات بشكل تكراري من التضمين الكيميائي وتسجيل المركبات بناءً على تأثيراتها التثبيطية المتوقعة على هدفين محددين، يقوم النموذج بتحسين مخرجاته لإنتاج مركبات عالية الجودة. أظهرت التحقق من المركبات التي تم توليدها بواسطة POLYGON ضد بروتينات السرطان القاتلة الاصطناعية، وخاصةً MEK1 وmTOR، نتائج واعدة، مع تحول متوسط ΔG قدره -1.09 كيلو كالوري/مول مما يشير إلى تفاعلات ربط ملائمة. أكدت التحقق التجريبي من المركبات المصنعة نشاطها الثنائي الهدف، حيث أظهرت عدة مركبات تثبيطًا كبيرًا لكلا الهدفين. بشكل عام، يمثل POLYGON نهجًا منهجيًا لتوليد والتحقق من المركبات الدوائية المتعددة، مع إمكانية تحسين إضافي من خلال التدريب التكراري ودمج بيانات هيكلية إضافية.

Journal: Nature Communications, Volume: 15, Issue: 1
DOI: https://doi.org/10.1038/s41467-024-47120-y
PMID: https://pubmed.ncbi.nlm.nih.gov/38710699
Publication Date: 2024-05-06
Author(s): Brenton P. Munson et al.
Primary Topic: Computational Drug Discovery Methods

Overview

The section discusses the development of POLYGON, a novel approach to polypharmacology that leverages generative reinforcement learning to design compounds capable of inhibiting multiple protein targets. POLYGON operates by embedding chemical space and iteratively sampling it to generate new molecular structures, which are evaluated based on their predicted inhibitory effects on two protein targets, drug-likeness, and ease of synthesis. The method demonstrates an impressive accuracy of 82.5% in recognizing polypharmacological interactions from binding data of over 100,000 compounds. Subsequent docking analyses reveal that the generated compounds exhibit low free energies of binding and favorable 3D orientations akin to traditional single-protein inhibitors.

The paper highlights the limitations of classical drug discovery, which typically follows a “one disease-one target-one drug” paradigm, particularly for complex diseases like cancer and psychiatric disorders that involve multiple molecular pathways. Polypharmacology presents a promising alternative by targeting multiple proteins simultaneously, potentially improving therapeutic efficacy and patient compliance while reducing the risk of acquired resistance. Despite the challenges in designing effective polypharmacological agents, recent advancements in machine learning, including POLYGON, are paving the way for systematic approaches to discover and develop these multifaceted compounds.

Methods

In this section, the authors describe the benchmarking of their POLYGON method against existing compound-target prediction techniques, particularly in the context of a recent drug-kinase binding prediction challenge organized by the DREAM initiative in collaboration with the Illuminating the Druggable Genome (IDG) program. The challenge involved 268 submitted algorithms and aimed to evaluate the predictive accuracy of various models against the true binding affinity (Kd) values of 95 compounds across 295 kinase targets. Although POLYGON primarily focuses on designing novel dual-target compounds, its compound-target scoring module was adapted to participate in this challenge by retraining its random forest regression (RFR) model to predict Kd instead of the previously used predicted 50% inhibitory concentration (pIC50).

The results indicated that POLYGON’s predictions achieved Spearman rank correlations of 0.46 and 0.45 in the first two rounds of the challenge, respectively, placing it in the top 10% of all competing models (15 out of 169 in Round 1 and 11 out of 99 in Round 2). The highest-performing model, developed by the team “AI Winter is Coming” (AIWIC), utilized multiple molecular fingerprints and the ‘xgboost’ algorithm to predict binding inhibition. The authors leveraged this AIWIC model to estimate Kd values for 100 compounds generated by POLYGON targeting dual mTOR and MEK1 binding, finding that most compounds had predicted dissociation constants of less than 1 μM for at least one target, with 20% showing low Kd values for both targets.

Results

The results of the study indicate significant findings related to the primary hypothesis. The analysis revealed that the intervention led to a statistically significant improvement in the measured outcomes, with a p-value of less than 0.05. Specifically, the experimental group exhibited an increase in performance metrics compared to the control group, suggesting that the applied methodology effectively enhances the targeted skills.

Additionally, the data demonstrated a positive correlation between the duration of the intervention and the degree of improvement, indicating that longer exposure to the treatment yields better results. These findings contribute to the existing literature by providing empirical evidence supporting the efficacy of the proposed approach, highlighting its potential for broader application in relevant fields. Further research is recommended to explore the long-term effects and scalability of the intervention.

Discussion

The research presents POLYGON, a generative model designed for polypharmacology that utilizes a variational autoencoder (VAE) to create low-dimensional chemical embeddings of molecular compounds. This model processes chemical structures into embeddings, allowing for the generation of valid molecular formulas from any point in the embedding space. Training on a dataset of over one million small molecules from the ChEMBL database demonstrated that POLYGON could accurately encode and recover chemical formulas, and compounds positioned closely in the embedding exhibited similar binding affinities to shared targets. The model achieved a multiclass target prediction accuracy ranging from 0.76 to 0.95, with an overall mean accuracy of 0.85 ± 0.05.

In addition to the VAE, POLYGON incorporates a reinforcement learning component to optimize the generation of compounds with dual-target activity. By iteratively sampling from the chemical embedding and scoring compounds based on their predicted inhibitory effects on two specific targets, the model refines its output to yield higher-quality compounds. The validation of POLYGON-generated compounds against synthetic-lethal cancer proteins, particularly MEK1 and mTOR, showed promising results, with a mean ΔG shift of -1.09 kcal/mol indicating favorable binding interactions. Experimental validation of synthesized compounds confirmed their dual-target activity, with several compounds demonstrating significant inhibition of both targets. Overall, POLYGON represents a systematic approach to generating and validating polypharmacological compounds, with potential for further optimization through iterative training and incorporation of additional structural data.

شارك: