بصمات جزيئية هي نماذج قوية لتوقع وظيفة الببتيد
Molecular fingerprints are strong models for peptide function prediction

شارك:
المجلة: Bioinformatics، المجلد: 42، العدد: 5
DOI: https://doi.org/10.1093/bioinformatics/btag179
PMID: https://pubmed.ncbi.nlm.nih.gov/41981726
تاريخ النشر: 2026-04-07
المؤلف: Jakub Adamczyk وآخرون
الموضوع الرئيسي: طرق اكتشاف الأدوية الحاسوبية

نظرة عامة

تبحث الدراسة في فعالية استخدام بصمات جزيئية بسيطة ومحددة المجال لتوقع خصائص الببتيدات، متchallenging الافتراض بأن نمذجة التفاعلات الجزيئية بعيدة المدى ضرورية. تستخدم الدراسة بصمات تعتمد على العد مثل بصمات الاتصال الموسعة (ECFP) والتواء الطوبولوجيا وبصمات RDKit بالتزامن مع LightGBM عبر 132 مجموعة بيانات، بما في ذلك LRGB وخمسة معايير إضافية للببتيدات. تكشف النتائج أن هذه النماذج، التي تركز على الميزات الجزيئية قصيرة المدى، تحقق دقة متقدمة، متفوقة على الشبكات العصبية الرسومية المعقدة (GNNs) والنهج المعتمد على المحولات. تؤكد التجارب الضابطة أن البصمات المحلية كافية للتنبؤات القوية، مما يثير تساؤلات حول ضرورة نمذجة التفاعلات بعيدة المدى.

في الختام، تقدم الدراسة نهجًا جديدًا لتوقع وظيفة الببتيد يستفيد من متغيرات العد لبصمات جزيئية مشفرة، مما يوضح كفاءتها الحاسوبية وقوتها التنبؤية دون الحاجة إلى معرفة هيكل الببتيد المطوي. يتجاوز المصنف القائم على ECFP بنصف قطر 2 النموذج الأفضل السابق في مجموعة بيانات Peptides-func بنسبة 1.5% AUPRC. يمثل هذا العمل المقارنة الأكثر شمولاً لبصمات الجزيئات المطبقة على الببتيدات حتى الآن، مما يبرز أهمية التصاميم المعتمدة على العد لتحقيق أداء قوي. ستستكشف الأبحاث المستقبلية تطبيق هذه المنهجية على البروتينات الأكبر والببتيدات المعدلة كيميائيًا، مما يعزز من إمكانية النماذج قصيرة المدى في مهام التعلم الجزيئي.

مقدمة

في المقدمة، يناقش المؤلفون أهمية الببتيدات، وهي سلاسل قصيرة من الأحماض الأمينية ذات وظائف بيولوجية متنوعة وإمكانات علاجية، خاصة كعوامل مضادة للميكروبات في ظل تزايد مقاومة الميكروبات. إن التنبؤ الدقيق بخصائص الببتيد أمر حاسم لاكتشاف الأدوية، ومع ذلك فإنه يقدم تحديات فريدة لتعلم الآلة (ML) بسبب قضايا مثل تحيز التشابه ومجموعات البيانات غير المتوازنة. لمعالجة هذه التحديات، تم تطوير مجموعات بيانات معيارية وإجراءات تقييم موحدة، وقد تحولت الاتجاهات الحديثة نحو استخدام نماذج لغة البروتين المدربة مسبقًا (PLMs) وتمثيلات الرسوم الجزيئية لنمذجة الببتيد.

يبرز المؤلفون قيود الأساليب الحالية، خاصة ندرة البيانات الهيكلية ثلاثية الأبعاد والتكاليف الحاسوبية المرتبطة بالحصول عليها. يقترحون إعادة النظر في بصمات الجزيئات المعتمدة على العد – الترميزات المدمجة المستمدة من تحت الرسوم الجزيئية الصغيرة – كميزات فعالة لتوقع وظيفة الببتيد. تظهر نتائجهم أنه عند دمجها مع مصنف LightGBM، تحقق هذه البصمات نتائج متقدمة عبر ستة معايير، بما في ذلك معيار الرسوم البيانية بعيدة المدى (LRGB)، دون الحاجة إلى ضبط معلمات فرعية أو معلومات هيكلية ثلاثية الأبعاد. يتحدى هذا الافتراض السائد بأن الاعتماديات بعيدة المدى حاسمة لتوقع وظيفة الببتيد، مما يشير إلى أن إحصائيات تحت الرسوم الجزيئية قصيرة المدى قد تكون كافية. يساهم المؤلفون من خلال (1) إظهار تفوق بصمات الجزيئات المعتمدة على العد على النماذج البديلة عبر 132 مجموعة بيانات، (2) تقديم معيار شامل لنماذج الببتيد المعتمدة على البصمات، و(3) تحليل متانة طريقتهم من خلال تجارب ضابطة.

طرق

في هذا القسم، يوضح المؤلفون بصمات الجزيئات المستخدمة في طريقتهم المقترحة لترميز الببتيد على مستوى الذرة. تهدف هذه الترميز إلى التقاط الخصائص الهيكلية والوظيفية المعقدة للببتيدات على المستوى الجزيئي. بالإضافة إلى ذلك، يوضح المؤلفون نهج معيار الرسوم البيانية بعيدة المدى، الذي يُستخدم لتقييم وجود الاعتماديات بعيدة المدى داخل هياكل الببتيد. تعتبر هذه المنهجية حاسمة لفهم كيف يمكن أن تؤثر التفاعلات البعيدة داخل الببتيد على سلوكه وخصائصه العامة.

نتائج

في هذا القسم، يقدم المؤلفون نتائج تجاربهم باستخدام نماذج تعتمد على البصمات المدربة على مجموعة متنوعة من المعايير التي تقيم خصائص مختلفة للببتيد. يلتزمون بالمنهجيات المعتمدة من الدراسات السابقة، ويبلغون عن مقاييس الأداء الرئيسية مثل المساحة تحت منحنى خاصية التشغيل (AUROC) ومعامل ارتباط ماثيوز (MCC)، مع توفر مقاييس إضافية في المواد التكميلية. تستخدم النماذج متغيرات العد وتُنفذ عبر مكتبة scikit-fingerprints، مع استخدام معلمات فرعية افتراضية: نصف قطر ECFP يبلغ 2 (قطر 4، ECFP4)، وطول مسار التواء الطوبولوجيا يبلغ 4، وطول مسار بصمة RDKit يبلغ 7، التي تركز على الاعتماديات قصيرة المدى.

يبرز المؤلفون أن اختيارهم لـ LightGBM مع 500 شجرة كمصنف يحقق أداءً قويًا، خاصة في سياق مجموعات البيانات غير المتوازنة، حيث يتم تطبيق وزن الفئة بشكل عكسي على تكرار الفئة الإيجابية. من الجدير بالذكر أنهم وجدوا أن ضبط المعلمات الفرعية عمومًا لم يؤدي إلى تحسينات كبيرة في الأداء، باستثناء حالات محددة مثل PeptideReactor. يتماشى التحسين المستمر للأداء مع زيادة عدد الأشجار مع الأدبيات الحالية. يتم التأكيد على الحساسية المنخفضة للطريقة المقترحة تجاه المعلمات الفرعية كميزة كبيرة، مما يسهل تقليل التكاليف الحاسوبية للت筛ين الافتراضي على نطاق واسع لعلاجات الببتيد وتأسيس خط أساس بسيط للأبحاث المستقبلية في هذا المجال.

نقاش

في قسم النقاش من ورقة البحث، يحلل المؤلفون أساليب التنبؤ المختلفة للببتيدات، مبرزين الاختلافات في طرق التمثيل. يصنفون هذه الطرق إلى أساليب قائمة على الرسوم البيانية على مستوى الذرة، التي تستخدم الشبكات العصبية الرسومية (GNNs) والمحولات الرسومية، وأنابيب قائمة على التسلسل تعتمد على الوصف الكيميائي الفيزيائي والمصنفات التقليدية. من الجدير بالذكر أن المؤلفين يبرزون فعالية نهجهم، الذي يستخدم بصمات جزيئية معتمدة على العد ومحددة المجال لترميز هيكل الببتيد. توفر هذه البصمات ميزات قوية وقليلة المعلمات تتفوق على نماذج التعلم العميق في مهام توقع وظيفة الببتيد، خاصة في السيناريوهات ذات البيانات المحدودة.

يستفيض المؤلفون في توضيح مزايا بصمات الجزيئات، التي تحول تصنيف الرسوم الجزيئية إلى تصنيف جدولي، مما يسمح بحساب فعال وقابل للتوسع. يجادلون بأن أداء النماذج المعتمدة على هذه البصمات يُعزى أساسًا إلى التفاعلات قصيرة المدى بدلاً من الاعتماديات بعيدة المدى، كما يتضح من نتائجهم المتفوقة على معيار الرسوم البيانية بعيدة المدى (LRGB). تتحدى النتائج الفكرة القائلة بأن التفاعلات بعيدة المدى ضرورية لتوقع وظيفة الببتيد بدقة، مما يشير بدلاً من ذلك إلى أن الهيكل الفطري للببتيدات يحد من أهمية مثل هذه التفاعلات. يخلص المؤلفون إلى أن نماذجهم المعتمدة على البصمات تحقق نتائج متقدمة عبر معايير مختلفة، مما يوضح متانتها وكفاءتها مقارنة بالنماذج الأكثر تعقيدًا التي تعتمد على التفاعلات بعيدة المدى.

Journal: Bioinformatics, Volume: 42, Issue: 5
DOI: https://doi.org/10.1093/bioinformatics/btag179
PMID: https://pubmed.ncbi.nlm.nih.gov/41981726
Publication Date: 2026-04-07
Author(s): Jakub Adamczyk et al.
Primary Topic: Computational Drug Discovery Methods

Overview

The research investigates the efficacy of using simple, domain-specific molecular fingerprints for peptide property prediction, challenging the assumption that modeling long-range molecular interactions is essential. The study employs count-based fingerprints such as Extended Connectivity Fingerprints (ECFP), Topological Torsion, and RDKit fingerprints in conjunction with LightGBM across 132 datasets, including LRGB and five additional peptide benchmarks. The findings reveal that these models, which focus on short-range molecular features, achieve state-of-the-art accuracy, outperforming more complex graph neural networks (GNNs) and transformer-based approaches. Control experiments validate that local fingerprints are sufficient for robust predictions, thereby questioning the necessity of long-range interaction modeling.

In conclusion, the study presents a novel approach to peptide function prediction that leverages count variants of hashed molecular fingerprints, demonstrating their computational efficiency and predictive power without requiring knowledge of the folded peptide structure. The ECFP-based classifier with a radius of 2 notably surpasses the previous best model on the Peptides-func dataset by 1.5% AUPRC. This work represents the most comprehensive comparison of molecular fingerprints applied to peptides to date, emphasizing the importance of count-based designs for achieving strong performance. Future research will explore the application of this methodology to larger proteins and chemically modified peptides, reinforcing the potential of short-range models in molecular learning tasks.

Introduction

In the introduction, the authors discuss the significance of peptides, which are short chains of amino acids with diverse biological functions and therapeutic potential, particularly as antimicrobial agents amid rising antimicrobial resistance. The accurate prediction of peptide properties is crucial for drug discovery, yet it presents unique challenges for machine learning (ML) due to issues like homology bias and imbalanced datasets. To address these challenges, various benchmark datasets and standardized evaluation procedures have been developed, and recent trends have shifted towards using pretrained protein language models (PLMs) and molecular graph representations for peptide modeling.

The authors highlight the limitations of existing approaches, particularly the scarcity of 3D structural data and the computational expense associated with obtaining it. They propose revisiting count-based molecular fingerprints—compact encodings derived from small molecular subgraphs—as effective features for peptide function prediction. Their findings demonstrate that when combined with a LightGBM classifier, these fingerprints achieve state-of-the-art results across six benchmarks, including the Long-Range Graph Benchmark (LRGB), without the need for hyperparameter tuning or 3D structural information. This challenges the prevailing assumption that long-range dependencies are critical for peptide function prediction, suggesting that short-range subgraph statistics may suffice. The authors contribute by (1) showing the superiority of count-based molecular fingerprints over alternative models across 132 datasets, (2) providing a comprehensive benchmark of fingerprint-based peptide models, and (3) analyzing the robustness of their method through controlled experiments.

Methods

In this section, the authors outline the molecular fingerprints utilized in their proposed atom-level peptide encoding method. This encoding aims to capture the intricate structural and functional characteristics of peptides at a molecular level. Additionally, the authors detail the Long Range Graph Benchmark approach, which is employed to assess the presence of long-range dependencies within the peptide structures. This methodology is crucial for understanding how distant interactions within the peptide may influence its overall behavior and properties.

Results

In this section, the authors present the results of their experiments using fingerprint-based models trained on a variety of benchmarks that assess different peptide properties. They adhere to established methodologies from previous studies, reporting key performance metrics such as the Area Under the Receiver Operating Characteristic curve (AUROC) and the Matthews Correlation Coefficient (MCC), with additional metrics available in the Supplementary Material. The models utilize count variants and are implemented via the scikit-fingerprints library, employing default hyperparameters: ECFP radius of 2 (diameter 4, ECFP4), Topological Torsion path length of 4, and RDKit fingerprint path length of 7, which focus on short-range dependencies.

The authors highlight that their choice of LightGBM with 500 trees as the classifier yields robust performance, particularly in the context of imbalanced datasets, where class weighting is applied inversely to the positive class frequency. Notably, they found that hyperparameter tuning generally did not lead to significant performance improvements, except in specific cases like PeptideReactor. The consistent enhancement of performance with an increased number of trees aligns with existing literature. The low sensitivity of the proposed method to hyperparameters is emphasized as a significant advantage, facilitating reduced computational costs for large-scale virtual screening of peptide therapeutics and establishing a straightforward baseline for future research in this domain.

Discussion

In the discussion section of the research paper, the authors analyze various prediction approaches for peptides, highlighting the differences in representation methods. They categorize these methods into atom-level graph-based approaches, which utilize message-passing Graph Neural Networks (GNNs) and Graph Transformers, and sequence-based pipelines that rely on physicochemical descriptors and classical classifiers. Notably, the authors emphasize the effectiveness of their approach, which employs domain-specific, count-based molecular fingerprints for peptide structure encoding. These fingerprints provide robust, low-parameter features that outperform deep learning models in peptide function prediction tasks, particularly in scenarios with limited data.

The authors further elaborate on the advantages of molecular fingerprints, which transform molecular graph classification into tabular classification, allowing for efficient computation and scalability. They argue that the performance of models based on these fingerprints is primarily attributed to short-range interactions rather than long-range dependencies, as evidenced by their superior results on the Long Range Graph Benchmark (LRGB). The findings challenge the notion that long-range interactions are essential for accurate peptide function prediction, suggesting instead that the inherent structure of peptides limits the relevance of such interactions. The authors conclude that their fingerprint-based models achieve state-of-the-art results across various benchmarks, demonstrating their robustness and efficiency compared to more complex models reliant on long-range interactions.

شارك: