تحليل واستكشاف شبكات الانتباه البيانية ونماذج اللغة المعتمدة على البروتين لتوقع تسلسلات بروتينات الطرد المقاومة لبورفيروموناس جينجيفاليس
Analyzing and exploring Graph Attention Networks and protein-based language models for predicting Porhyromonas gingivalis resistant efflux protein sequences

شارك:
المجلة: Dental and Medical Problems، المجلد: 62، العدد: 2
DOI: https://doi.org/10.17219/dmp/186143
PMID: https://pubmed.ncbi.nlm.nih.gov/40372388
تاريخ النشر: 2025-04-30
المؤلف: Pradeep Kumar Yadalam وآخرون
الموضوع الرئيسي: الميكروبيولوجيا الفموية وبحوث التهاب اللثة

نظرة عامة

تتناول ورقة البحث القضية الحرجة لمقاومة المضادات الحيوية (AMR) من خلال الاستفادة من نماذج حسابية متقدمة للتنبؤ بآليات المقاومة في *Porphyromonas gingivalis*. من خلال التركيز على الجينومات ذات الأولوية العالية لمقاومة المضادات الحيوية، تهدف الدراسة إلى تعزيز استراتيجيات العلاج والحد من تفشي العدوى المقاومة للمضادات الحيوية على مستوى العالم. استخدم المؤلفون تسلسلات بروتينات الطرد المتعددة المقاومة من *P. gingivalis*، التي تم تحديدها بشكل خاص من خلال UniProt ID A0A0K2J2N6_PORGN، وقاموا بتحليل هذه التسلسلات باستخدام إطار عمل DeepBIO، الذي يجمع بين نماذج اللغة الكبيرة (LLMs) والشبكات العميقة ذات الانتباه.

أظهرت النتائج أن نماذج الذاكرة طويلة وقصيرة المدى (LSTM)-الانتباه، ProtBERT، وBERTGAT حققت درجات حساسية عالية بلغت 0.9 ومعدلات دقة بلغت 89.5%، 88.5%، و90.5%، على التوالي، في تحديد السلالات المقاومة. بالإضافة إلى ذلك، كانت درجات الخصوصية لهذه النماذج 0.89، 0.87، و0.90، مما يدل على فعاليتها في تحديد الحالات غير المقاومة بدقة، وهو أمر أساسي لتقليل الإيجابيات الكاذبة في الكشف عن AMR. تشير النتائج إلى أن تطبيق LLMs في البيئات السريرية يمكن أن يساهم بشكل كبير في منع AMR من خلال تمكين استراتيجيات العلاج المستهدفة التي تتناول جينات المقاومة المحددة، وبالتالي التخفيف من التهديد العالمي الذي تشكله مقاومة المضادات الحيوية.

مقدمة

تتناول مقدمة ورقة البحث القضية الحرجة لمقاومة المضادات الحيوية (AMR)، والتي تُعرف بأنها قدرة الكائنات الدقيقة على مقاومة تأثيرات الأدوية المضادة للميكروبات، بما في ذلك المضادات الحيوية. يؤكد المؤلفون على ضرورة التنبؤ بـ AMR لمكافحة الأمراض المقاومة للمضادات الحيوية بشكل فعال، خاصة من خلال مراقبة الجينومات ذات الأولوية العالية لمقاومة المضادات الحيوية. يقترحون الاستفادة من نماذج اللغة الكبيرة (LLMs)، مثل ProtBERT وBERTGAT، للتحقيق في العلاقات المعقدة بين تسلسلات البروتين، وتكويناتها الهيكلية، ووظائفها البيولوجية. من خلال اعتبار تسلسلات البروتين ككيانات لغوية، تهدف هذه المقاربة إلى توضيح الروابط المعقدة بين ترتيبات الأحماض الأمينية وهياكلها ثلاثية الأبعاد الناتجة، والتي تعتبر أساسية لفهم العمليات البيولوجية.

تسلط الورقة الضوء على أهمية *Porphyromonas gingivalis*، وهي بكتيريا مرتبطة بالتهاب اللثة، كدراسة حالة لمقاومة المضادات الحيوية. إن زيادة مقاومة *P. gingivalis* للمضادات الحيوية، والتي تعود بشكل كبير إلى عمل مضخات الطرد المقاومة (RND)، تؤكد على ضرورة فهم آليات مقاومتها. يشير المؤلفون إلى أن بروتينات مثل AcrA وAcrB وTolC تلعب أدوارًا حاسمة في هذه المقاومة المتعددة الأدوية (MDR). بالإضافة إلى ذلك، تحدد الدراسة إنزيم إصلاح NAD(P)H-الهيدرات ثنائي الوظيفة A0A212GBI3_PORGN كبروتين طرد مقاوم شائع. تهدف الأبحاث إلى استخدام شبكات الانتباه البيانية (GATs) ونماذج اللغة المعتمدة على البروتين للتنبؤ بتسلسلات بروتينات الطرد المقاومة في *P. gingivalis*، مما يساهم في استراتيجيات التخفيف من AMR.

طرق البحث

في هذه الدراسة، استخدم المؤلفون قاعدة بيانات UniProt لتنزيل تسلسلات البروتينات المقاومة المتعددة من *Porphyromonas gingivalis*. بشكل خاص، ركزوا على إدخالات البروتين التالية: A0A0K2J2N6_PORGN وA0A212GBI3_PORGN. إن اختيار هذه التسلسلات أمر حاسم لفهم آليات مقاومة الأدوية في هذا الممرض، وهو أمر مهم في سياق مرض التهاب اللثة وعلاجه. تضمن المنهج المستخدم أن يكون التحليل مستندًا إلى بيانات بروتين موثوقة وشاملة، مما يسهل المزيد من التحقيقات في الآثار الوظيفية لهذه البروتينات المقاومة.

النتائج

في هذه الدراسة، تم استخدام نماذج LSTMattention وProtBERT وBERTGAT لاستخراج الميزات المخفية والأوزان من تسلسلات البروتين FASTA. تم ضبط النماذج باستخدام خوارزميات الانتشار العكسي مع مُحسِّن ADAM على مدى 50 تكرارًا. وُجد أن حساسية النماذج كانت 0.90 لـ LSTMattention وProtBERT، و0.91 لـ BERTGAT، مما يدل على معدل إيجابي حقيقي مرتفع (TP / (TP + FN)).

بالإضافة إلى ذلك، تم الإبلاغ عن الخصوصية، التي تقيس معدل السلبيات الحقيقية، كـ 0.89 لـ LSTMattention، و0.87 لـ ProtBERT، و0.90 لـ BERTGAT، مما يعكس نسبة السلبيات الفعلية التي تم تحديدها بدقة (TN / (TN + FP)). تشير هذه النتائج إلى أن النماذج الثلاثة تظهر أداءً قويًا في كل من الحساسية والخصوصية عند تحليل تسلسلات البروتين.

المناقشة

تسلط المناقشة الضوء على فعالية نماذج الذكاء الاصطناعي المتقدمة، وخاصة LSTMattention وProtBERT وBERTGAT، في التنبؤ بمقاومة المضادات الحيوية (AMR) في *Porphyromonas gingivalis*، حيث حققت معدلات دقة تصل إلى 90.5%. من بين هذه النماذج، أظهرت BERTGAT الأداء الأعلى نظرًا لتضمينها آليات الانتباه المعتمدة على الرسوم البيانية، مما يعزز قدرة النموذج على التقاط العلاقات البيولوجية المعقدة. تؤكد الدراسة على أهمية استهداف بروتينات الطرد في *P. gingivalis* لتصميم أدوية مضادة حيوية جديدة، نظرًا لدورها في مقاومة المضادات الحيوية. تم التحقق من قابلية تفسير النماذج وقوتها من خلال مجموعة متنوعة من الطرق التحليلية، بما في ذلك SHAP وUMAP ومخططات UpSet، مما يشير إلى تطبيقات سريرية محتملة للكشف عن السلالات المقاومة.

على الرغم من هذه النتائج الواعدة، فإن البحث محدود من خلال مجموعة بيانات صغيرة وغياب التحقق الخارجي، مما يشير إلى الحاجة إلى مزيد من الدراسات مع مجموعات بيانات أكبر وأكثر تنوعًا لتعزيز موثوقية وعمومية النماذج التنبؤية. تؤكد المناقشة على الدور الحاسم لتعلم الآلة في معالجة AMR في التهاب اللثة، مما يبرز ضرورة البحث المستمر لاستكشاف قابلية تطبيق هذه النماذج عبر مختلف مسببات الأمراض الفموية.

Journal: Dental and Medical Problems, Volume: 62, Issue: 2
DOI: https://doi.org/10.17219/dmp/186143
PMID: https://pubmed.ncbi.nlm.nih.gov/40372388
Publication Date: 2025-04-30
Author(s): Pradeep Kumar Yadalam et al.
Primary Topic: Oral microbiology and periodontitis research

Overview

The research paper addresses the critical issue of antimicrobial resistance (AMR) by leveraging advanced computational models to predict resistance mechanisms in *Porphyromonas gingivalis*. By focusing on high-priority AMR genomes, the study aims to enhance treatment strategies and curb global outbreaks of antibiotic-resistant infections. The authors utilized multi-drug-resistant efflux protein sequences from *P. gingivalis*, specifically identified through UniProt ID A0A0K2J2N6_PORGN, and analyzed these sequences using the DeepBIO framework, which combines large language models (LLMs) with deep attention networks.

The results demonstrated that the Long Short-Term Memory (LSTM)-attention, ProtBERT, and BERTGAT models achieved high sensitivity scores of 0.9 and accuracy rates of 89.5%, 88.5%, and 90.5%, respectively, in identifying resistant strains. Additionally, the specificity scores for these models were 0.89, 0.87, and 0.90, indicating their effectiveness in accurately identifying non-resistant cases, which is essential for reducing false positives in AMR detection. The findings suggest that the application of LLMs in clinical settings could significantly contribute to preventing AMR by enabling targeted treatment strategies that address specific resistance genes, thereby mitigating the global threat posed by antibiotic resistance.

Introduction

The introduction of the research paper addresses the critical issue of antimicrobial resistance (AMR), defined as the ability of microorganisms to withstand the effects of antimicrobial drugs, including antibiotics. The authors emphasize the necessity of predicting AMR to effectively combat antibiotic-resistant diseases, particularly through the surveillance of high-priority AMR genomes. They propose leveraging large language models (LLMs), such as ProtBERT and BERTGAT, to investigate the complex relationships between protein sequences, their structural configurations, and biological functions. By treating protein sequences as linguistic entities, this approach aims to elucidate the intricate connections between amino acid arrangements and their resultant three-dimensional structures, which are fundamental to understanding biological processes.

The paper highlights the significance of Porphyromonas gingivalis, a bacterium implicated in periodontitis, as a case study for AMR. The increasing resistance of P. gingivalis to antibiotics, largely due to the action of resistance-nodulation division (RND) efflux pumps, underscores the urgency of understanding its resistance mechanisms. The authors note that proteins such as AcrA, AcrB, and TolC play crucial roles in this multidrug resistance (MDR). Additionally, the study identifies the bifunctional NAD(P)H-hydrate repair enzyme A0A212GBI3_PORGN as a prevalent resistant efflux protein. The research aims to utilize Graph Attention Networks (GATs) and protein-based language models to predict resistant efflux protein sequences in P. gingivalis, thereby contributing to strategies for mitigating AMR.

Methods

In this study, the authors utilized the UniProt database to download sequences of multidrug-resistant proteins from *Porphyromonas gingivalis*. Specifically, they focused on the following protein entries: A0A0K2J2N6_PORGN and A0A212GBI3_PORGN. The selection of these sequences is critical for understanding the mechanisms of drug resistance in this pathogen, which is significant in the context of periodontal disease and its treatment. The methodology employed ensures that the analysis is grounded in reliable and comprehensive protein data, facilitating further investigation into the functional implications of these resistance proteins.

Results

In this study, the models LSTMattention, ProtBERT, and BERTGAT were employed to extract hidden features and weights from FASTA protein sequences. The models were fine-tuned using backpropagation algorithms with the ADAM optimizer over 50 iterations. The sensitivity of the models was found to be 0.90 for LSTMattention and ProtBERT, and 0.91 for BERTGAT, indicating a high true positive rate (TP / (TP + FN)).

Additionally, the specificity, which measures the true negative rate, was reported as 0.89 for LSTMattention, 0.87 for ProtBERT, and 0.90 for BERTGAT, reflecting the proportion of actual negatives accurately identified (TN / (TN + FP)). These results suggest that all three models demonstrate strong performance in both sensitivity and specificity when analyzing protein sequences.

Discussion

The discussion highlights the efficacy of advanced artificial intelligence models, specifically LSTMattention, ProtBERT, and BERTGAT, in predicting antimicrobial resistance (AMR) in *Porphyromonas gingivalis*, achieving accuracy rates up to 90.5%. Among these, BERTGAT demonstrated the highest performance due to its incorporation of graph-based attention mechanisms, which enhance the model’s ability to capture complex biological relationships. The study underscores the importance of targeting efflux proteins in *P. gingivalis* for novel antibiotic drug design, given their role in antibiotic resistance. The interpretability and robustness of the models were validated through various analytical methods, including SHAP, UMAP, and UpSet plots, suggesting potential clinical applications for detecting resistant strains.

Despite these promising findings, the research is limited by a small dataset and the absence of external validation, indicating a need for further studies with larger and more diverse datasets to enhance the reliability and generalizability of the predictive models. The discussion emphasizes the critical role of machine learning in addressing AMR in periodontitis, highlighting the necessity for ongoing research to explore the applicability of these models across different oral pathogens.

شارك: