DOI: https://doi.org/10.1186/s12864-025-11600-2
PMID: https://pubmed.ncbi.nlm.nih.gov/40269681
تاريخ النشر: 2025-04-23
المؤلف: Gerold Csendes وآخرون
الموضوع الرئيسي: علم النسخ الجيني أحادي الخلية والمكاني
نظرة عامة
في هذا القسم، يؤكد المؤلفون على الحاجة الملحة للتنبؤات الدقيقة للاستجابات الخلوية للت perturbations لتعزيز فهمنا لسلوك الخلايا في حالات مختلفة. يشيرون إلى ندرة بيانات الاضطراب مقارنة ببيانات الخلايا الأساسية، مما أدى إلى تطوير نماذج خلوية أساسية تستخدم بيانات التعبير الجيني على مستوى الخلية الواحدة على نطاق واسع. هذه النماذج، بما في ذلك scGPT و scFoundation، تم ضبطها بدقة لمهام تنبؤية محددة ولكنها تواجه تحديات في القياس المناسب.
تقوم الدراسة بتقييم scGPT و scFoundation مقابل النماذج الأساسية وتكشف عن نتائج مفاجئة: حتى أبسط نموذج أساسي، الذي يقوم بمتوسط أمثلة التدريب، تفوق على كلا النموذجين الأساسيين. علاوة على ذلك، تفوقت نماذج التعلم الآلي الأساسية التي تستفيد من الميزات البيولوجية ذات الصلة بشكل كبير على scGPT. كما يشير المؤلفون إلى أن مجموعات بيانات Perturb-Seq الحالية تظهر تباينًا منخفضًا خاصًا بالاضطرابات، مما يجعلها غير كافية لتقييم النموذج. تؤكد هذه النتائج على قيود منهجيات القياس الحالية وتقترح الحاجة إلى استراتيجيات محسنة في تقييم نماذج التنبؤ بتعبير الجينات بعد الاضطراب.
مقدمة
في مقدمة هذه الورقة البحثية، يتناول المؤلفون تعقيدات نمذجة الفينوتيب الخلوي، وهو أمر حاسم لفهم كل من الحالات الصحية والمريضة. يسلطون الضوء على تطور النماذج الحاسوبية من المنطق البولياني والمعادلات التفاضلية العادية إلى الأساليب الحديثة للتعلم العميق، مع التركيز بشكل خاص على فائدة بيانات التعبير الجيني المستمدة من تقنيات التسلسل المختلفة. تعتبر بيانات التعبير الجيني بمثابة بديل قيم للحالات الخلوية، خاصة عند تحليل بيانات ما بعد الاضطراب، التي تعكس العلاقات السببية بين الاضطرابات وتغيرات التعبير الجيني.
يناقش المؤلفون ظهور نماذج أساسية تعتمد على المحولات، مثل scGPT و scFoundation، والتي تم تدريبها مسبقًا على بيانات تسلسل RNA أحادية الخلية واسعة النطاق. تهدف هذه النماذج إلى تعزيز التنبؤات بالفينوتيب بعد الاضطراب. ومع ذلك، تكشف الدراسة أن هذه النماذج الأساسية غالبًا ما تؤدي أداءً أقل مقارنة بالنماذج الأساسية الأبسط التي تستخدم متوسط عينات التدريب. بالإضافة إلى ذلك، تتفوق نماذج التعلم الآلي التقليدية التي تتضمن معرفة بيولوجية سابقة بشكل كبير على النماذج الأساسية. كما يشير المؤلفون إلى أن التباين المنخفض بين العينات في مجموعات البيانات الحالية يشكل تحديات لتقييم أداء النموذج بدقة، خاصة في سيناريوهات القياس التي تركز على سياقات اضطراب مختلفة.
طرق
في هذه الدراسة، قام المؤلفون بتقييم نموذجين من نماذج اللغة الكبيرة (LLM) المعتمدة على بنى المحولات، scGPT و scFoundation، لتنبؤ RNA-seq بعد الاضطراب. تم تدريب كلا النموذجين مسبقًا على بيانات تسلسل RNA أحادية الخلية غير المصنفة واسعة النطاق لتعلم تمثيلات الجينات وعلاقات الجينات. تتنبأ النماذج بملفات RNA-seq للخلايا المضطربة من خلال استخدام متجهات RNA-seq من الخلايا غير المضطربة وتمثيل الاضطراب. يتضمن scGPT رمز اضطراب لنمذجة التأثيرات، بينما يستخدم scFoundation تمثيلات الجينات المدربة مسبقًا ضمن إطار شبكة عصبية رسومية للتنبؤات. تم ضبط النماذج بدقة على مجموعات بيانات القياس التي تم إنشاؤها من خلال Perturb-seq، والتي تدمج الاضطرابات المعتمدة على CRISPR مع التسلسل أحادي الخلية.
ركز تقييم أداء النموذج على قدرتهم على التعامل مع الاضطرابات غير المرئية، باستخدام مقاييس مثل معاملات الارتباط بيرسون في كل من المساحات التعبيرية الخام والتفاضلية. أشارت النتائج إلى أن حتى أبسط نموذج أساسي، Train Mean، تفوق على النماذج الأساسية في التنبؤ بملفات التعبير الجيني بعد الاضطراب. على وجه التحديد، تفوق نموذج الانحدار العشوائي (RF) مع ميزات علم الأحياء (GO) بشكل كبير على أداء كل من scGPT و scFoundation عبر مجموعات بيانات متعددة. تشير النتائج إلى أن النماذج الأساسية واجهت صعوبة في تعلم تمثيلات ذات معنى بيولوجي للاضطرابات، كما يتضح من أدائها المنخفض مقارنة بالنماذج الأبسط. بالإضافة إلى ذلك، سلطت الدراسة الضوء على أهمية تحليل التعبير التفاضلي، خاصة في سياق الجينات المستهدفة في تجارب CRISPR، والتي أثرت على قدرات النماذج التنبؤية. بشكل عام، تؤكد النتائج فعالية النماذج التنبؤية الأبسط مقارنة بالنماذج الأساسية الأكثر تعقيدًا في هذا السياق.
النتائج
يقدم قسم “النتائج” من الورقة البحثية النتائج الرئيسية المستمدة من التجارب أو التحليلات التي تم إجراؤها. يوضح نتائج الاختبارات المختلفة، مع تسليط الضوء على الاتجاهات والأنماط المهمة التي لوحظت في البيانات. غالبًا ما تكون النتائج مصحوبة بتحليلات إحصائية، بما في ذلك قيم p وفترات الثقة، للتحقق من النتائج.
بالإضافة إلى ذلك، قد يتضمن القسم تمثيلات رسومية، مثل المخططات أو الجداول، لتوضيح العلاقات بين المتغيرات أو فعالية التدخلات. تعزز هذه المساعدات البصرية وضوح النتائج وتسهيل المقارنة عبر ظروف أو مجموعات مختلفة. بشكل عام، تسهم النتائج في فهم سؤال البحث وقد تشير إلى تداعيات للدراسات المستقبلية أو التطبيقات العملية.
المناقشة
في هذه المناقشة، يحلل المؤلفون أداء نماذج مختلفة، بما في ذلك نموذج Train Mean البسيط، مقابل نماذج أساسية أكثر تعقيدًا مثل scGPT و scFoundation في التنبؤ بنتائج تسلسل RNA بعد الاضطراب باستخدام مجموعات بيانات Perturb-seq. بشكل مفاجئ، غالبًا ما تفوق نموذج Train Mean على هذه النماذج المتقدمة، بينما تفوق نموذج الانحدار العشوائي الذي دمج المعرفة البيولوجية السابقة عبر مصطلحات علم الأحياء (GO) بشكل كبير عليها. تشير هذه النتيجة إلى أن مجموعات بيانات القياس، وخاصة دراسات Adamson و Replogle RPE، قد تفتقر إلى التباين الكافي لتمييز أداء النماذج بشكل فعال. يبرز المؤلفون أنه على الرغم من أن هذه المجموعات تحتوي على عدد كبير من الخلايا الفردية، إلا أن العدد المحدود من الاضطرابات المتميزة يؤدي إلى تباين بيولوجي منخفض، مما يقوض القدرة على تقييم النماذج الأكثر تعقيدًا بدقة.
تكشف التحليلات أيضًا عن علاقة عكسية بين الوسيط بيرسون لملفات الاضطراب وأداء النماذج، مما يشير إلى أن مجموعات البيانات ذات التباين الأكبر تكون أكثر فعالية في القياس. يجادل المؤلفون بأن التركيز الحالي على مشكلة PEX، التي تتنبأ بالاستجابات للاضطرابات الجديدة، يتجاهل مشكلة CEX التي تتضمن التنبؤات في أنواع خلايا جديدة. يقترحون أن مجموعات البيانات التي تحتوي على مجموعة أوسع من الاضطرابات، مثل LINCS-L1000، يمكن أن تعزز القياسات المستقبلية. بالإضافة إلى ذلك، تثير الدراسة مخاوف بشأن فائدة بيانات تسلسل RNA أحادية الخلية للتنبؤات بعد الاضطراب، حيث أدت النماذج الأساسية التي تستخدم بيانات شبه جماعية إلى أداء مماثل أو أفضل من النماذج الأساسية. بشكل عام، يدعو المؤلفون إلى مزيد من القياسات الدقيقة التي تتضمن تباينًا أعلى ومجموعات بيانات متنوعة لتقييم نماذج التعلم الآلي بشكل أفضل في مهام التنبؤ بعد الاضطراب.
DOI: https://doi.org/10.1186/s12864-025-11600-2
PMID: https://pubmed.ncbi.nlm.nih.gov/40269681
Publication Date: 2025-04-23
Author(s): Gerold Csendes et al.
Primary Topic: Single-cell and spatial transcriptomics
Overview
In this section, the authors emphasize the critical need for accurate predictions of cellular responses to perturbations to enhance our understanding of cell behavior in various states. They note the scarcity of perturbation data compared to baseline cellular data, which has led to the development of foundation cell models utilizing large-scale single-cell gene expression data. These models, including scGPT and scFoundation, are fine-tuned for specific predictive tasks but face challenges in proper benchmarking.
The study benchmarks scGPT and scFoundation against baseline models and reveals surprising findings: even the simplest baseline model, which averages training examples, outperformed both foundation models. Moreover, basic machine learning models that leverage biologically relevant features significantly surpassed scGPT. The authors also point out that existing Perturb-Seq benchmark datasets show low variance specific to perturbations, rendering them inadequate for model evaluation. These findings underscore the limitations of current benchmarking methodologies and suggest the need for improved strategies in assessing post-perturbation gene expression prediction models.
Introduction
In the introduction of this research paper, the authors address the complexities of modeling cellular phenotypes, which is crucial for understanding both healthy and diseased states. They highlight the evolution of computational models from Boolean logic and ordinary differential equations to contemporary deep learning approaches, particularly emphasizing the utility of transcriptomics data derived from various sequencing techniques. Transcriptomics serves as a valuable proxy for cellular states, especially when analyzing post-perturbation data, which reflects the causal relationships between perturbations and gene expression changes.
The authors discuss the emergence of Transformer-based foundation models, such as scGPT and scFoundation, which are pre-trained on extensive single-cell RNA sequencing data. These models aim to enhance predictions of post-perturbation phenotypes. However, the study reveals that these foundation models often underperform compared to simpler baseline models that utilize the mean of training samples. Additionally, traditional machine learning models that incorporate biological prior knowledge significantly outperform foundation models. The authors also note that the low inter-sample variance in existing datasets poses challenges for accurately assessing model performance, particularly in benchmarking scenarios that focus on different perturbation contexts.
Methods
In this study, the authors benchmarked two large language model (LLM)-based transformer architectures, scGPT and scFoundation, for post-perturbation RNA-seq prediction. Both models were pre-trained on extensive unlabelled single-cell RNA sequencing (scRNA-seq) data to learn gene embeddings and gene-gene relationships. The models predict RNA-seq profiles of perturbed cells by utilizing RNA-seq vectors from unperturbed cells and a representation of the perturbation. scGPT incorporates a perturbation token to model effects, while scFoundation employs pretrained gene embeddings within a graph neural network framework for predictions. The models were fine-tuned on benchmark datasets generated through Perturb-seq, which integrates CRISPR-based perturbations with single-cell sequencing.
The evaluation of model performance focused on their ability to handle unseen perturbations, utilizing metrics such as Pearson correlation coefficients in both raw and differential expression spaces. Results indicated that even the simplest baseline model, Train Mean, outperformed the foundation models in predicting post-perturbation gene expression profiles. Specifically, the Random Forest Regressor (RF) with Gene Ontology (GO) features significantly surpassed the performance of both scGPT and scFoundation across multiple datasets. The findings suggest that the foundation models struggled to learn biologically meaningful representations of perturbations, as evidenced by their lower performance compared to simpler models. Additionally, the study highlighted the importance of differential expression analysis, particularly in the context of target genes in CRISPR experiments, which influenced the models’ predictive capabilities. Overall, the results underscore the effectiveness of simpler predictive models over more complex foundation models in this context.
Results
The “Results” section of the research paper presents key findings derived from the conducted experiments or analyses. It details the outcomes of various tests, highlighting significant trends and patterns observed in the data. The results are often accompanied by statistical analyses, including p-values and confidence intervals, to validate the findings.
Additionally, the section may include graphical representations, such as charts or tables, to illustrate the relationships between variables or the effectiveness of interventions. These visual aids enhance the clarity of the results and facilitate comparison across different conditions or groups. Overall, the findings contribute to the understanding of the research question and may suggest implications for future studies or practical applications.
Discussion
In this discussion, the authors analyze the performance of various models, including a simple Train Mean model, against more complex foundation models like scGPT and scFoundation in predicting post-perturbation RNA sequencing outcomes using Perturb-seq datasets. Surprisingly, the Train Mean model often outperformed these advanced models, while a Random Forest Regressor that incorporated biological prior knowledge via Gene Ontology (GO) terms significantly surpassed them. This finding suggests that the benchmark datasets, particularly the Adamson and Replogle RPE studies, may lack sufficient variability to effectively differentiate model performances. The authors highlight that while these datasets contain a large number of single cells, the limited number of distinct perturbations results in low biological variance, which undermines the ability to assess more complex models accurately.
The analysis further reveals an inverse relationship between the median Pearson correlation of perturbation profiles and the performance of models, indicating that datasets with greater heterogeneity are more effective for benchmarking. The authors argue that the current focus on the PEX problem, which predicts responses to novel perturbations, neglects the CEX problem involving predictions in novel cell types. They suggest that datasets with a broader range of perturbations, such as LINCS-L1000, could enhance future benchmarks. Additionally, the study raises concerns about the utility of single-cell RNA-seq data for post-perturbation predictions, as baseline models using pseudo-bulk data performed comparably or better than foundation models. Overall, the authors call for more rigorous benchmarks that incorporate higher variance and diverse datasets to better evaluate machine learning models in post-perturbation prediction tasks.
