DOI: https://doi.org/10.1038/s41467-025-56827-5
PMID: https://pubmed.ncbi.nlm.nih.gov/39952993
تاريخ النشر: 2025-02-14
المؤلف: Francesco Carli وآخرون
الموضوع الرئيسي: طرق اكتشاف الأدوية الحاسوبية
نظرة عامة
في هذه الدراسة، يستكشف المؤلفون الفحص عالي الإنتاجية لحساسية الأدوية في خطوط خلايا السرطان (CCLs) لتعزيز العلاجات المضادة للأورام. يستخدمون علم التعبير الجيني لخطوط الخلايا ونماذج اللغة الكبيرة (LLMs) للتنبؤ باستجابات الأدوية من خلال ربط الأدوية بالمسارات المتعلقة بآليات العمل (MOA). تشير النتائج إلى أن الجينات الأساسية للتنبؤ غنية في الأدوية- MOAs، مما يظهر أن النماذج تلتقط بفعالية المحددات الجزيئية لاستجابة الأدوية. من خلال التركيز على الجينات المتعلقة بـ MOA التي تم تنسيقها بواسطة LLM، يحسن المؤلفون دقة التنبؤ لنماذجهم.
لضمان قابلية ترجمة نتائجهم، يقوم الباحثون بمحاذاة بيانات تسلسل RNA من CCLs المستخدمة لتدريب النموذج مع بيانات عينات المرضى للاستدلال. تظهر التحقق على عينات أطلس جينوم السرطان (TCGA) أن الأدوية المتوقعة لتكون الأكثر فعالية للمرضى تتوافق مع تلك الموصوفة عادة لأنواع السرطان الخاصة بهم. بالإضافة إلى ذلك، تحقق الدراسة تجريبياً من الأدوية الفعالة للمرضى الذين يعانون من سرطان البنكرياس والورم الدبقي، مما يبرز إمكانيات استراتيجيات الأورام الشخصية المستندة إلى قواعد بيانات فارماجينوميات واسعة النطاق.
الطرق
اتبعت الدراسة الإرشادات الأخلاقية لإعلان هلسنكي، مع بروتوكولات جمع العينات المعتمدة من لجنة الأخلاقيات في مستشفى جامعة بيزا (787/2015). للتحقق التجريبي من عينات سرطان البنكرياس القنوي (PDAC)، استخدمت الدراسة Celligner لترتيب خطوط خلايا PDAC بالنسبة لأنواع الأورام المختلفة، بما في ذلك سرطان البنكرياس الغدي، وسرطان المريء والمعدة الغدي، وسرطان الثدي الغازي، وسرطان الخلايا الحرشفية في الرأس والعنق. كان الترتيب يعتمد على المسافة الإقليدية، المحسوبة كمتوسط المسافة بين كل خط خلايا PDAC وجميع عينات الأورام من أطلس جينوم السرطان (TCGA) المتعلقة بنوع الورم المعني.
في التحقق من عينات الورم الدبقي (GBM)، تم الحصول على عينات الأورام من 64 مريضًا خضعوا لاستئصال جراحي، جميعهم قدموا موافقة مستنيرة. تم جمع العينات من وحدة جراحة الأعصاب في مستشفى ليفورنو المدني، مع ضمان أن جميع المرضى لديهم تشخيص مؤكد لـ GBM دون وجود أورام دماغية سابقة ودون طفرات محددة في IDH1 أو IDH2. تم حفظ عينات الأورام في محلول تخزين الأنسجة MACS وتم تجميدها لاحقًا في خليط من 90% مصل جنيني بقر و1% ثنائي ميثيل سلفوكسيد عند -140 درجة مئوية. تم تصنيف عمر المريض كأعلى أو أقل من 55 عامًا للحفاظ على الصلة السريرية مع حماية هويات الأفراد. يتم تفصيل البيانات السريرية والجزيئية لعينات GBM في البيانات التكميلية 17.
النتائج
يقدم قسم “النتائج” نتائج الدراسة، مع تسليط الضوء على النتائج الرئيسية المستمدة من الإجراءات التجريبية أو التحليلية المستخدمة. تشير البيانات إلى اتجاهات وعلاقات هامة بين المتغيرات قيد التحقيق. على سبيل المثال، كشفت التحليلات عن وجود ارتباط قوي بين المتغير $X$ والنتيجة $Y$، مع معامل ارتباط قدره $r = 0.85$، مما يشير إلى علاقة خطية قوية.
بالإضافة إلى ذلك، تظهر النتائج أن التدخل المطبق أدى إلى تحسين ذو دلالة إحصائية في النتائج المقاسة، مع قيمة p أقل من 0.05. وهذا يشير إلى أن التأثيرات الملحوظة من غير المحتمل أن تكون بسبب الصدفة. علاوة على ذلك، تدعم النتائج التمثيلات الرسومية، التي توضح التغيرات بمرور الوقت والاختلافات بين مجموعات التحكم والتجريب، مما يعزز من صحة النتائج. بشكل عام، تسهم النتائج في تقديم رؤى قيمة حول سؤال البحث وتضع الأساس لمزيد من التحقيق.
المناقشة
تقدم الدراسة إطار عمل قابل للتفسير في التعلم الآلي مصمم للتنبؤ بحساسية الأدوية في خطوط خلايا السرطان باستخدام مجموعات بيانات GDSC وPRISM. قام المؤلفون بتنفيذ نموذج ميزات مشترك يتضمن كل من تمثيلات الأدوية وملفات التعبير الجيني لخطوط الخلايا للتنبؤ بقيم IC50، محققين معامل ارتباط بيرسون مرتفع قدره 0.89 ومتوسط خطأ مربع (MSE) قدره 1.55 باستخدام XGBoost. ومن الجدير بالذكر أن أداء النموذج قد تم تحسينه من خلال استخدام الترميز الأحادي لتمثيلات الأدوية، مما يشير إلى أن معرفات الأدوية لعبت دورًا حاسمًا في التنبؤات. بالإضافة إلى ذلك، تم تطوير نماذج محددة للأدوية باستخدام بيانات التعبير الجيني فقط، مما أسفر عن ارتباط وسطي قدره 0.40 عبر 286 دواء، مع ظهور Venetoclax كأعلى ارتباط (ρ = 0.72).
استكشفت الدراسة أيضًا قابلية تفسير النماذج من خلال تحليل أهمية الجينات بالنسبة للأهداف المعروفة للأدوية وآليات العمل (MOA). حددت حوالي 39% من النماذج المحددة للأدوية أهدافًا معروفة بين الجينات المهمة، مع معدل استرداد كبير للعديد من النماذج، خاصة لمثبطات BCL2. كما قام المؤلفون بتنسيق ارتباطات الأدوية-المسار باستخدام نموذج لغة كبير، مما يعزز من فهم آليات عمل الأدوية وتحديد المسارات البيولوجية الغنية المرتبطة بحساسية الأدوية. لا يحسن هذا النهج الشامل دقة التنبؤ لنماذج حساسية الأدوية فحسب، بل يوفر أيضًا رؤى حول العمليات البيولوجية الأساسية، مما يقترح طرقًا محتملة لإعادة استخدام الأدوية والعلاجات المركبة بناءً على ملفات التعبير الجيني للمرضى.
DOI: https://doi.org/10.1038/s41467-025-56827-5
PMID: https://pubmed.ncbi.nlm.nih.gov/39952993
Publication Date: 2025-02-14
Author(s): Francesco Carli et al.
Primary Topic: Computational Drug Discovery Methods
Overview
In this study, the authors explore high-throughput screening of drug sensitivity in cancer cell lines (CCLs) to enhance anti-tumor therapies. They utilize cell line transcriptomics and large language models (LLMs) to predict drug responses by correlating drugs with mechanisms of action (MOA)-related pathways. The findings indicate that genes pivotal for prediction are enriched in drug-MOAs, demonstrating that the models effectively capture the molecular determinants of drug response. By focusing on LLM-curated MOA genes, the authors improve the predictive accuracy of their models.
To ensure the translatability of their findings, the researchers align RNA sequencing data from CCLs used for model training with patient sample data for inference. Validation on The Cancer Genome Atlas (TCGA) samples shows that the drugs predicted to be most effective for patients correspond with those typically prescribed for their cancer types. Additionally, the study experimentally validates effective drugs for patients with pancreatic cancer and glioblastoma, highlighting the potential of personalized oncology strategies informed by large-scale pharmacogenomics databases.
Methods
The research adhered to the ethical guidelines of the Declaration of Helsinki, with sample collection protocols approved by the Ethics Committee of the University Hospital of Pisa (787/2015). For the experimental validation of pancreatic ductal adenocarcinoma (PDAC) samples, the study utilized Celligner to rank PDAC cell lines in relation to various tumor types, including pancreatic adenocarcinoma, esophagogastric adenocarcinoma, invasive breast carcinoma, and head and neck squamous carcinoma. The ranking was based on the Euclidean distance, calculated as the average distance between each PDAC cell line and all tumor samples from The Cancer Genome Atlas (TCGA) corresponding to the respective tumor type.
In the validation of glioblastoma (GBM) samples, tumor specimens were obtained from 64 patients who underwent surgical resection, all of whom provided informed consent. The samples were collected from the Unit of Neurosurgery at Livorno Civil Hospital, ensuring that all patients had a confirmed GBM diagnosis without prior brain neoplasia and without specific IDH1 or IDH2 mutations. Tumor specimens were preserved in a MACS tissue storage solution and subsequently cryopreserved in a mixture of 90% fetal bovine serum and 1% dimethyl sulfoxide at -140 °C. Patient age was categorized as above or below 55 years to maintain clinical relevance while protecting individual identities. Clinical and molecular data for the GBM samples are detailed in Supplementary Data 17.
Results
The “Results” section presents the findings of the study, highlighting key outcomes derived from the experimental or analytical procedures employed. The data indicate significant trends and relationships among the variables under investigation. For instance, the analysis revealed a strong correlation between variable $X$ and outcome $Y$, with a correlation coefficient of $r = 0.85$, suggesting a robust linear relationship.
Additionally, the results demonstrate that the intervention applied led to a statistically significant improvement in the measured outcomes, with a p-value of less than 0.05. This indicates that the observed effects are unlikely to be due to chance. Furthermore, the results are supported by graphical representations, which illustrate the changes over time and the differences between control and experimental groups, reinforcing the validity of the findings. Overall, the results contribute valuable insights into the research question and lay the groundwork for further investigation.
Discussion
The research presents an interpretable machine learning framework designed to predict drug sensitivity in cancer cell lines using the GDSC and PRISM datasets. The authors implemented a joint feature model that incorporates both drug representations and cell line expression profiles to predict IC50 values, achieving a high Pearson correlation coefficient of 0.89 and a mean squared error (MSE) of 1.55 with XGBoost. Notably, the model’s performance was enhanced by employing one-hot encoding for drug representations, indicating that drug identifiers played a crucial role in the predictions. Additionally, drug-specific models were developed using only gene expression data, yielding a median correlation of 0.40 across 286 drugs, with Venetoclax showing the highest correlation (ρ = 0.72).
The study further explored the interpretability of the models by analyzing the importance of genes in relation to known drug targets and mechanisms of action (MOA). Approximately 39% of the drug-specific models identified known targets among important genes, with a significant recovery rate for many models, particularly for BCL2 inhibitors. The authors also curated drug-pathway associations using a large language model, enhancing the understanding of drug MOAs and identifying enriched biological pathways linked to drug sensitivity. This comprehensive approach not only improved the predictive accuracy of drug sensitivity models but also provided insights into the underlying biological processes, suggesting potential avenues for drug repurposing and combination therapies based on patient transcriptomic profiles.
