DOI: https://doi.org/10.1038/s41598-024-62419-y
PMID: https://pubmed.ncbi.nlm.nih.gov/38796582
تاريخ النشر: 2024-05-25
المؤلف: Mariana del Carmen Aguilera‐Puga وآخرون
الموضوع الرئيسي: الببتيدات المضادة للميكروبات والأنشطة
نظرة عامة
تبحث هذه الدراسة في تطبيق نماذج التعلم الآلي في اكتشاف وتصميم الببتيدات النشطة حيوياً، مع التركيز بشكل خاص على الببتيدات المضادة للميكروبات المصنفة على أنها مدمرة للغشاء، مخترقة للغشاء، ومرتبطة بالبروتينات. تؤكد الدراسة على أهمية الوعي بهيكل البروتين في هذه النماذج، التي تعتمد تقليدياً على البيانات التسلسلية ولكنها غالباً ما تكافح لالتقاط الآليات المعقدة للعمل. من خلال تحليل ميزات مثل الديببتيدات والخصائص الفيزيائية الكيميائية، حقق المؤلفون دقة تنبؤية عالية (86-88%) لهذه الفئات من الببتيدات. ومع ذلك، أظهرت النماذج الأولية تحيزاً نحو الهياكل الحلزونية α والهياكل الملفوفة، مما دفع إلى تنفيذ استراتيجيات اختيار المجموعات الفرعية وتقليل البيانات للتخفيف من هذا التحيز.
تكشف النتائج أن النماذج الخاصة بالهيكل، التي تم تدريبها على مجموعات فرعية من الببتيدات التي من المحتمل أن تتبنى هياكل معينة، تفوقت بشكل عام على النماذج غير المعتمدة على الهيكل. ومن الجدير بالذكر أن أداء النماذج تدهور عندما تمت إزالة التسلسلات الحلزونية α بشكل عشوائي، مما يبرز ضرورة التوازن الدقيق في البيانات لمنع فقدان المعلومات الحيوية. تؤكد الدراسة على حساسية الميزات الرئيسية عبر فئات هيكلية مختلفة وتدعو إلى مراعاة الفروق الهيكلية في تصنيف الببتيدات. تمثل هذه الدراسة جهدًا رائدًا لمعالجة التحيز الهيكلي في النماذج التنبؤية، مما يبرز أهمية علاقات التسلسل-الهيكل-الوظيفة في تصميم العلاجات المعتمدة على الببتيدات من خلال الذكاء الاصطناعي.
الطرق
في قسم الطرق، يحدد المؤلفون سير عمل شامل لتوقع النشاط الغشائي أو البروتيني لتسلسلات الببتيد، مع دمج كل من المعلومات المستندة إلى التسلسل والمعلومات الهيكلية. جمعوا 1,057 تسلسل ببتيد من قواعد بيانات عامة متنوعة واستخرجوا 8,437 ميزة تتعلق بتراكيب الأحماض الأمينية والخصائص الفيزيائية الكيميائية العالمية. بعد تصفية البيانات من أجل الأهمية، قاموا ببناء نماذج تعلم آلي باستخدام 12 خوارزمية للتصنيف الثنائي و9 للتصنيف الثلاثي، بهدف التمييز بين الببتيدات المدمرة للغشاء (MDPs)، والببتيدات القابلة للاختراق عبر الغشاء (MPPs)، والأدوية المعتمدة على الببتيدات (PBPs). خضعت النماذج للتحسين من خلال التحقق المتقاطع بعشرة أضعاف، وتقنيات زيادة العينة، وضبط المعلمات، تلتها اختبار ضد مجموعة تحقق خارجية.
لمعالجة الطبيعة غير المتوازنة لمجموعات البيانات المتعلقة بفئات الببتيدات الثلاث، استخدم المؤلفون ثلاث طرق لزيادة العينة: SMOTE (تقنية زيادة العينة للأقليات الاصطناعية)، ROSE (أمثلة عشوائية لزيادة العينة)، وADASYN (العينة الاصطناعية التكيفية). بالإضافة إلى ذلك، قاموا بالتحقيق في الخصائص الهيكلية لمجموعاتهم، مشيرين إلى وجود تحيز كبير نحو الحلزونات α. وقد دفع ذلك إلى مراجعة مجموعات التدريب للتركيز إما على هذه الفئة الهيكلية المحددة أو لإنشاء نموذج أكثر عمومية من خلال توليد تسلسلات شبيهة بالأصل من فئات هيكلية أقل انتشارًا.
النتائج
يقدم قسم “النتائج” من ورقة البحث النتائج الرئيسية المستمدة من التجارب أو التحليلات التي تم إجراؤها. عادةً ما يتضمن بيانات كمية، وتحليلات إحصائية، وتمثيلات بصرية مثل الرسوم البيانية أو الجداول التي توضح النتائج. غالبًا ما تتم مقارنة النتائج مع الفرضيات أو الدراسات السابقة لتسليط الضوء على الفروقات أو التأكيدات المهمة.
في هذا القسم، قد يبلغ المؤلفون عن مقاييس محددة، مثل المتوسطات، والانحرافات المعيارية، أو قيم p، لدعم ادعاءاتهم. بالإضافة إلى ذلك، يتم مناقشة أي اتجاهات أو أنماط ملحوظة في البيانات، مما يوفر رؤى حول تداعيات النتائج. بشكل عام، تخدم النتائج في التحقق من أهداف البحث وتساهم في الفهم الأوسع للموضوع قيد التحقيق.
المناقشة
في هذا القسم، يوضح المؤلفون منهجيتهم في تنسيق وتحليل تسلسلات الببتيد من قواعد بيانات متنوعة، مع التركيز على تصنيفها إلى فئات متميزة بناءً على آليات العمل: الببتيدات المدمرة للغشاء/المخترقة للغشاء (MDPs)، الببتيدات المخترقة للغشاء (MPPs)، والببتيدات المرتبطة بالبروتينات (PBPs). تم جمع ما مجموعه 1,057 ببتيد، مع تحديد 415 MDPs، 334 MPPs، و308 PBPs. استخدمت الدراسة عملية شاملة لاستخراج الميزات، حيث تم قياس تركيب الأحماض الأمينية، وتركيبات الديببتيدات والترايببتيدات، والخصائص الفيزيائية الكيميائية، مما أسفر عن 8,537 خاصية مستمدة من التسلسل. لتعزيز أداء النموذج، تم معالجة التعدد الخطي من خلال إزالة الميزات المرتبطة بشكل كبير، مما أدى في النهاية إلى الاحتفاظ بـ 53 إلى 57 خاصية ذات صلة لمهام التصنيف.
نفذ المؤلفون خوارزميات تعلم آلي متنوعة لتصنيف الببتيدات، حيث أظهر مصنف الغابة العشوائية (RFC) أداءً متفوقًا عبر مهام التصنيف الثنائي والثلاثي. استخدمت الدراسة التحقق المتقاطع بعشرة أضعاف وتقنيات زيادة العينة للتخفيف من عدم توازن الفئات، محققة دقة تنبؤية عالية (حتى 88.0% للتصنيفات الثنائية و86.7% للتصنيفات الثلاثية). كشفت تحليل التنوع الهيكلي عن هيمنة الهياكل الحلزونية α بين الببتيدات، مما يشير إلى أن نماذج التصنيف قد تكون فعالة بشكل خاص للببتيدات التي من المحتمل أن تتبنى مثل هذه التكوينات. بشكل عام، تؤكد النتائج على أهمية الخصائص الفيزيائية الكيميائية المحددة في التمييز بين فئات الببتيدات وتبرز فعالية نماذج التعلم الآلي المعتمدة على الأشجار في مهام تصنيف الببتيدات.
DOI: https://doi.org/10.1038/s41598-024-62419-y
PMID: https://pubmed.ncbi.nlm.nih.gov/38796582
Publication Date: 2024-05-25
Author(s): Mariana del Carmen Aguilera‐Puga et al.
Primary Topic: Antimicrobial Peptides and Activities
Overview
This research investigates the application of machine learning models in the discovery and design of bioactive peptides, specifically focusing on antimicrobial peptides categorized as membrane-disrupting, membrane-penetrating, and protein-binding peptides. The study emphasizes the importance of protein structure awareness in these models, which traditionally rely on sequential data but often struggle to capture the complex mechanisms of action. By analyzing features such as dipeptides and physicochemical properties, the authors achieved high predictive accuracy (86-88%) for these peptide categories. However, initial models exhibited a bias towards α-helical and coiled structures, prompting the implementation of subset selection and data reduction strategies to mitigate this bias.
The findings reveal that structure-specific models, which were trained on subsets of peptides likely to adopt specific structures, generally outperformed structure-agnostic models. Notably, the performance of models deteriorated when α-helical sequences were removed indiscriminately, highlighting the necessity for careful data balancing to prevent loss of critical information. The research underscores the sensitivity of key features across different structural classes and advocates for the consideration of structural nuances in peptide classification. This study represents a pioneering effort to address structural bias in predictive models, emphasizing the relevance of sequence-structure-function relationships in the design of peptide-based therapeutics through artificial intelligence.
Methods
In the Methods section, the authors outline a comprehensive workflow for predicting the membrane or protein activity of peptide sequences, incorporating both sequence-based and structural information. They collected 1,057 peptide sequences from various public databases and extracted 8,437 features related to amino acid compositions and global physicochemical properties. After filtering for significance, they constructed machine learning models using 12 algorithms for binary classification and 9 for ternary classification, aimed at distinguishing between membrane-disrupting peptides (MDPs), membrane-permeable peptides (MPPs), and peptide-based drugs (PBPs). The models underwent optimization through tenfold cross-validation, oversampling techniques, and hyperparameter tuning, followed by testing against an external validation set.
To address the imbalanced nature of the datasets corresponding to the three peptide classes, the authors employed three oversampling methods: SMOTE (Synthetic Minority Oversampling Technique), ROSE (Random Over-Sampling Examples), and ADASYN (Adaptive Synthetic Sampling). Additionally, they investigated the structural characteristics of their datasets, noting a significant bias towards α-helices. This prompted a revision of the training sets to either focus on this specific structural class or to create a more generalized model by generating native-like sequences from less prevalent structural classes.
Results
The “Results” section of the research paper presents the key findings derived from the conducted experiments or analyses. It typically includes quantitative data, statistical analyses, and visual representations such as graphs or tables that illustrate the outcomes. The results are often compared against the hypotheses or previous studies to highlight significant differences or confirmations.
In this section, the authors may report specific metrics, such as means, standard deviations, or p-values, to substantiate their claims. Additionally, any observed trends or patterns in the data are discussed, providing insights into the implications of the findings. Overall, the results serve to validate the research objectives and contribute to the broader understanding of the topic under investigation.
Discussion
In this section, the authors detail their methodology for curating and analyzing peptide sequences from various databases, focusing on their classification into distinct categories based on mechanisms of action: membranolytic/membrane-disrupting peptides (MDPs), membrane-penetrating peptides (MPPs), and protein-binding peptides (PBPs). A total of 1,057 peptides were collected, with 415 MDPs, 334 MPPs, and 308 PBPs identified. The study employed a comprehensive feature extraction process, measuring amino acid composition, dipeptide and tripeptide compositions, and physicochemical properties, resulting in 8,537 sequence-derived properties. To enhance model performance, multicollinearity was addressed by eliminating highly correlated features, ultimately retaining 53 to 57 relevant properties for classification tasks.
The authors implemented various machine learning algorithms to classify the peptides, with Random Forest Classifier (RFC) demonstrating superior performance across both binary and ternary classification tasks. The study utilized tenfold cross-validation and oversampling techniques to mitigate class imbalances, achieving high prediction accuracies (up to 88.0% for binary and 86.7% for ternary classifications). The analysis of structural diversity revealed a predominance of α-helical structures among the peptides, suggesting that the classification models may be particularly effective for peptides likely to adopt such conformations. Overall, the findings underscore the importance of specific physicochemical properties in distinguishing between peptide classes and highlight the efficacy of tree-based machine learning models in peptide classification tasks.
