تحليل الذكاء لكفاءة توصيل جزيئات الأدوية إلى مواقع الأورام السرطانية باستخدام نماذج التعلم الآلي
Intelligence analysis of drug nanoparticles delivery efficiency to cancer tumor sites using machine learning models

شارك:
المجلة: Scientific Reports، المجلد: 15، العدد: 1
DOI: https://doi.org/10.1038/s41598-024-84450-9
PMID: https://pubmed.ncbi.nlm.nih.gov/39762427
تاريخ النشر: 2025-01-06
المؤلف: Wael A. Mahdi وآخرون
الموضوع الرئيسي: طرق اكتشاف الأدوية الحاسوبية

نظرة عامة

تستكشف هذه الدراسة تطبيق نماذج التعلم الآلي (ML) للتنبؤ بتوزيع الجزيئات النانوية في مختلف الأعضاء، باستخدام مجموعة بيانات تركز على سلوك الجزيئات النانوية في علاج السرطان. تشمل مجموعة البيانات متغيرات فئوية ورقمية تتعلق بخصائص الجزيئات النانوية، تستهدف بشكل خاص توزيعها في أعضاء مثل الورم، القلب، الكبد، الطحال، الرئة، والكلى. لمعالجة الطبيعة المعقدة وغير الخطية للبيانات بشكل فعال، استخدم الباحثون ثلاثة نماذج تعلم آلي: الانحدار البايزي (BRR)، والانحدار القائم على النواة (KRR)، وأقرب الجيران (KNN). تم اختيار هذه النماذج بناءً على قدراتها في إدارة تعقيد البيانات والعلاقات غير الخطية.

لتحسين أداء النموذج، دمجت الدراسة تقنيات متقدمة مثل خوارزمية اليراعة لضبط المعلمات الفائقة وإزالة الميزات التكرارية (RFE) لاختيار الميزات. أشارت النتائج إلى أن KRR تفوقت على النماذج الأخرى، محققة درجات R² أعلى وقيم جذر متوسط مربع الخطأ (RMSE) أقل لمعظم معلمات الإخراج. تؤكد النتائج فعالية KRR في التقاط الخصائص غير الخطية لتوزيع الجزيئات النانوية بدقة، مما يوفر رؤى قيمة لتحسين النماذج التنبؤية في هذا المجال.

النتائج

في هذا القسم، تم تقييم أداء ثلاثة نماذج تعلم آلي—الانحدار البايزي (BRR)، والانحدار القائم على النواة (KRR)، وأقرب الجيران (KNN)—لتنبؤ توزيع الجزيئات النانوية في مختلف الأعضاء، بما في ذلك الورم، القلب، الكبد، الطحال، الرئة، والكلى. تم تقييم النماذج باستخدام درجات R² المعتمدة على التحقق المتقاطع ومقاييس جذر متوسط مربع الخطأ (RMSE). تفوقت KRR على النماذج الأخرى، محققة أعلى درجات R² وأقل قيم RMSE، مما يدل على فعاليتها في التقاط الخصائص غير الخطية لمجموعة البيانات. على سبيل المثال، حققت KRR درجة R² اختبارية قدرها 0.703 وRMSE قدره 2.605 لإخراج الورم DE، متجاوزة بشكل كبير درجة R² للنموذج الأساسي البالغة 0.41 وRMSE البالغة 2.02.

بالإضافة إلى ذلك، شمل التحليل عملية اختيار الميزات، التي حددت الميزات الرئيسية التي تسهم في دقة التنبؤ لكل عضو. تسرد الجدول 4 هذه الميزات، مثل الحجم، والجهد الزتاوي، ومواد الجزيئات النانوية المختلفة، جنبًا إلى جنب مع متوسط نسبة التحسين في قيم R² بعد اختيار الميزات. أشارت النتائج إلى أن KRR تفوقت باستمرار على النموذج الأساسي عبر عدة مخرجات، خاصة بالنسبة للقلب DE (R² = 0.900، RMSE = 0.631) والطحال DE (R² = 0.964، RMSE = 0.897)، على الرغم من أنها لم تتجاوز النموذج الأساسي بالنسبة للرئة DE. بشكل عام، تؤكد النتائج قوة نموذج KRR في التنبؤ بتوزيع الجزيئات النانوية، مع تقديم تمثيلات بصرية لدقة النموذج في الأشكال المرفقة.

المناقشة

في هذا القسم، يناقش المؤلفون المنهجية المستخدمة لإعداد البيانات والنمذجة في دراستهم حول توزيع الجزيئات النانوية عبر مختلف الأعضاء. تتكون مجموعة البيانات، المشتقة من دراسة بحثية سابقة، من 534 صفًا مع متغيرات فئوية (مثل النوع، MAT، TS) ورقمية (مثل الحجم، الجهد الزتاوي، الإدارة)، والتي تعتبر ضرورية لتقدير كفاءة التوزيع (DE) في أعضاء مثل الورم، القلب، الكبد، الطحال، الرئة، والكلى. استخدم المؤلفون الاستيفاء الخطي لملء القيم المفقودة، والترميز الأحادي للميزات الفئوية، وطريقة Z-score لاكتشاف القيم الشاذة، مما يضمن سلامة وجودة البيانات. بالإضافة إلى ذلك، تم استخدام MaxAbsScaler لتطبيع الميزات، مما يعزز أداء نماذج التعلم الآلي من خلال الحفاظ على تشتت البيانات.

شملت مرحلة النمذجة تطبيق ثلاثة خوارزميات تعلم آلي: الانحدار البايزي (BRR)، والانحدار القائم على النواة (KRR)، وأقرب الجيران (KNN). تم استخدام خوارزمية اليراعة لضبط المعلمات الفائقة، بينما تم تنفيذ إزالة الميزات التكرارية (RFE) لاختيار الميزات، مما يحسن أداء النموذج. أشارت النتائج إلى أن KRR تفوقت على النماذج الأخرى في تقدير نتائج توزيع الجزيئات النانوية، محققة أعلى درجات R² وأقل قيم RMSE. تؤكد هذه الدراسة فعالية تقنيات التعلم الآلي المتقدمة في التقاط العلاقات المعقدة وغير الخطية المتأصلة في بيانات توزيع الجزيئات النانوية.

Journal: Scientific Reports, Volume: 15, Issue: 1
DOI: https://doi.org/10.1038/s41598-024-84450-9
PMID: https://pubmed.ncbi.nlm.nih.gov/39762427
Publication Date: 2025-01-06
Author(s): Wael A. Mahdi et al.
Primary Topic: Computational Drug Discovery Methods

Overview

This study investigates the application of machine learning (ML) models to predict the biodistribution of nanoparticles in various organs, utilizing a dataset focused on nanoparticle behavior in cancer treatment. The dataset encompasses both categorical and numerical variables related to nanoparticle properties, specifically targeting their distribution in organs such as the tumor, heart, liver, spleen, lung, and kidney. To effectively address the complex and non-linear nature of the data, the researchers employed three ML models: Bayesian Ridge Regression (BRR), Kernel Ridge Regression (KRR), and K-Nearest Neighbors (KNN). The selection of these models was informed by their capabilities in managing data complexity and non-linear relationships.

To enhance model performance, the study incorporated advanced techniques such as the Firefly Algorithm for hyperparameter tuning and Recursive Feature Elimination (RFE) for feature selection. The results indicated that KRR outperformed the other models, achieving higher R² scores and lower root mean square error (RMSE) values for most output parameters. The findings underscore the efficacy of KRR in accurately capturing the non-linear characteristics of nanoparticle biodistribution, providing valuable insights for optimizing predictive models in this domain.

Results

In this section, the performance of three machine learning models—Bayesian Ridge Regression (BRR), Kernel Ridge Regression (KRR), and K-Nearest Neighbors (KNN)—was evaluated for predicting the biodistribution of nanoparticles in various organs, including tumor, heart, liver, spleen, lung, and kidney. The models were assessed using cross-validated R² scores and root mean square error (RMSE) metrics. KRR outperformed the other models, achieving the highest R² scores and the lowest RMSE values, thus demonstrating its effectiveness in capturing the dataset’s nonlinear characteristics. For instance, KRR achieved a test R² of 0.703 and an RMSE of 2.605 for the DE tumor output, significantly surpassing the baseline model’s R² of 0.41 and RMSE of 2.02.

Additionally, the analysis included a feature selection process, which identified key features contributing to the predictive accuracy for each organ. Table 4 lists these features, such as size, zeta potential, and various nanoparticle materials, along with the mean percentage improvement in R² values post-feature selection. The results indicated that KRR consistently outperformed the baseline model across multiple outputs, particularly for DE heart (R² = 0.900, RMSE = 0.631) and DE spleen (R² = 0.964, RMSE = 0.897), although it did not surpass the baseline for DE lung. Overall, the findings underscore the robustness of the KRR model in predicting nanoparticle biodistribution, with visual representations of model accuracy provided in the accompanying figures.

Discussion

In this section, the authors discuss the methodology employed for data preparation and modeling in their study on the biodistribution of nanoparticles across various organs. The dataset, derived from a prior research study, comprises 534 rows with both categorical (e.g., Type, MAT, TS) and numerical variables (e.g., Size, Zeta Potential, Admin), which are crucial for estimating distribution efficiency (DE) in organs such as the tumor, heart, liver, spleen, lung, and kidney. The authors utilized linear interpolation for missing value imputation, one-hot encoding for categorical features, and the Z-score method for outlier detection, ensuring data integrity and quality. Additionally, MaxAbsScaler was employed for feature normalization, enhancing the performance of machine learning models by maintaining data sparsity.

The modeling phase involved the application of three machine learning algorithms: Bayesian Ridge Regression (BRR), Kernel Ridge Regression (KRR), and K-nearest neighbors (KNN). The Firefly Algorithm was utilized for hyperparameter tuning, while Recursive Feature Elimination (RFE) was implemented for feature selection, optimizing model performance. The results indicated that KRR outperformed the other models in estimating biodistribution outcomes, achieving the highest R² scores and lowest RMSE values. This study underscores the effectiveness of advanced machine learning techniques in capturing the complex, non-linear relationships inherent in nanoparticle biodistribution data.

شارك: