تقييم منهجي لطرق الذكاء الاصطناعي القابلة للتفسير لتوقع بقاء مرضى السرطان بناءً على النسخ الجينية عالية الأبعاد
A systematic evaluation of explainable AI methods for high-dimensional transcriptome-based cancer survival prediction

شارك:
المجلة: Frontiers in Physiology، المجلد: 17
DOI: https://doi.org/10.3389/fphys.2026.1830956
PMID: https://pubmed.ncbi.nlm.nih.gov/42099920
تاريخ النشر: 2026-04-22
المؤلف: 左熠熠 وآخرون
الموضوع الرئيسي: الذكاء الاصطناعي القابل للتفسير (XAI)

نظرة عامة

تناقش هذه القسم دور الذكاء الاصطناعي القابل للتفسير (XAI) في معالجة التحديات التي تطرحها طبيعة “الصندوق الأسود” لنماذج التعلم العميق في توقع بقاء مرضى السرطان بناءً على بيانات النسخ الجيني. يبرز المؤلفون غياب المعايير النظامية لتقييم أطر XAI المصممة خصيصًا لبيانات البقاء عالية الأبعاد. لمعالجة ذلك، قاموا بتقييم ست طرق تمثيلية لـ XAI – مصنفة كطرق قائمة على التدرج، وطرق قائمة على الانتشار، وطرق قائمة على الاضطراب – باستخدام شبكة عصبية ذاتية التطبيع (SNN) كنموذج أساسي. استخدمت التحليلات 6,248 عينة من 15 نوعًا من السرطان مأخوذة من أطلس جينوم السرطان (TCGA) ضمن إطار موحد.

ركزت مقاييس التقييم على ثلاثة أبعاد: إثراء العوامل التنبؤية (المقاسة من خلال دلالة الانحدار الأحادي لكوكس)، التناسق البيولوجي (تم التحقق منه ضد أربعة قواعد بيانات موثوقة، بما في ذلك OpenTargets)، واستقرار التفسير (تم تقييمه باستخدام مؤشر كونشيفا). كشفت النتائج أن DeepSHAP تفوقت على الطرق الأخرى، حيث حددت أكثر العوامل التنبؤية دلالة إحصائية مع ضمان استقرار عالٍ في التفسير. أظهرت LRP (انتشار الأهمية على مستوى الطبقات) خصوصية تنبؤية أقل قليلاً لكنها تفوقت في التوافق مع قواعد البيانات البيولوجية، مما جعلها فعالة في التحقق من الصلة البيولوجية. على العكس، أظهرت الطريقة القائمة على الاضطراب، PFI (أهمية ميزة التبديل)، قيودًا كبيرة، خاصة في التعامل مع التداخل بين الميزات، مما أدى إلى استقرار منخفض. تختتم الدراسة بتأسيس استقرار التفسير كمؤشر موثوق للصلاحية البيولوجية في XAI، مما يوفر إطارًا لاختيار أدوات تفسير الذكاء الاصطناعي الفعالة في الطب الدقيق.

الطرق

في هذا القسم، يوضح المؤلفون تصميمهم التجريبي والإطار الإحصائي لتقييم أداء طرق الذكاء الاصطناعي القابل للتفسير (XAI) في تحديد العوامل التنبؤية من بيانات النسخ الجيني للسرطان الشامل. استخدموا استراتيجية التحقق المتقاطع المتكرر 10× 5-fold، مما أسفر عن 50 نموذجًا مستقلًا لكل مجموعة سرطان. تم استخدام مجموعات التدريب والتحقق الداخلية لمراقبة التوقف المبكر، وتم إجراء 100 عينة Bootstrap لتقدير فترات الثقة لتحليلات الارتباط. وجدت الدراسة أن DeepSHAP تفوقت بشكل كبير على طرق XAI الأخرى، مثل PFI و LRP، في تحديد الجينات التنبؤية ذات الصلة سريريًا عبر 15 نوعًا من السرطان، حيث حققت DeepSHAP وسطيًا قدره 37.5 جينًا ذا دلالة مقارنة بتقليصات قدرها 23.7% و 31.6% لـ PFI و LRP، على التوالي (اختبار ويلكوكسون، p المعدل وفقًا لمعدل الخطأ الكاذب < 0.001). كما كشفت التحليلات أن الطرق القائمة على التدرج والانتشار أظهرت استقرارًا أكبر في نسبة الميزات مقارنة بالطرق القائمة على الاضطراب، كما يتضح من درجات مؤشر كونشيفا الأعلى (اختبار فريدمان، p < 0.001). احتلت DeepSHAP باستمرار مرتبة بين أفضل الطرق عبر أنواع السرطان المختلفة، لا سيما في الورم الدبقي وورم الدبقيات منخفضة الدرجة، حيث حددت أكبر عدد من العوامل التنبؤية مع توزيع مضغوط عبر النماذج. علاوة على ذلك، لوحظت علاقة إيجابية قوية بين أداء النموذج التنبؤي (مؤشر C) واستقرار التفسير، لا سيما في المجموعات التي تتمتع بقدرات تنبؤية قوية. وهذا يشير إلى أن نسب الميزات المستقرة من المرجح أن تعكس إشارات تنبؤية ذات معنى بيولوجي، بينما قد تلتقط الطرق غير المستقرة مثل PFI تباينات عشوائية بدلاً من رؤى بيولوجية هامة.

النتائج

يقدم قسم “النتائج” من ورقة البحث النتائج المستمدة من البرنامج المنفذ والبيئة التجريبية. تشير النتائج الرئيسية إلى فعالية المنهجية المقترحة في تحقيق الأهداف المرجوة. تظهر النتائج تحسينات كبيرة في مقاييس الأداء مقارنة بالنماذج الأساسية، مما يبرز قوة النهج تحت ظروف مختلفة.

تؤكد التحليلات الإحصائية موثوقية النتائج، حيث تشير قيم p إلى دلالة قوية. بالإضافة إلى ذلك، توضح التمثيلات المرئية للبيانات، مثل الرسوم البيانية والمخططات، المزايا التنافسية للطريقة الجديدة، مما يعزز قابليتها للتطبيق في السيناريوهات الواقعية. بشكل عام، تسهم النتائج في تقديم رؤى قيمة في هذا المجال وتقترح طرقًا للبحث المستقبلي.

المناقشة

في هذه الدراسة، قمنا بتحليل مجموعة بيانات شاملة للسرطان تضم 6,248 عينة من المرضى عبر 15 نوعًا من السرطان، باستخدام إطار توقع البقاء العميق القائم على وحدات الخطية الأسية المقاسة (SELU) وAlpha Dropout لمعالجة تحديات بيانات النسخ الجيني عالية الأبعاد. حقق نموذجنا متوسط مؤشر التوافق (C-index) قدره 0.622، مما يدل على أداء قوي في توقع البقاء. استخدمنا عدة طرق للذكاء الاصطناعي القابل للتفسير (XAI)، بما في ذلك التدرجات المتكاملة، DeepLIFT، وانتشار الأهمية على مستوى الطبقات (LRP)، لتقييم قابلية تفسير نموذجنا. ومن الجدير بالذكر أن LRP أظهرت أعلى تناسق بيولوجي مع قواعد بيانات جينات السرطان المعتمدة، بينما برزت DeepSHAP كأكثر الطرق فعالية في تحديد العلامات البيولوجية التنبؤية، مما يوازن بين إثراء التنبؤ، التناسق البيولوجي، واستقرار التفسير.

تشير نتائجنا إلى أن الطرق القائمة على الانتشار، لا سيما DeepSHAP، تتفوق في السياقات الجينومية عالية الأبعاد، حيث إنها تتجنب إنتاج مدخلات غير قابلة للتفسير بيولوجيًا، على عكس الطرق القائمة على الاضطراب مثل أهمية ميزة التبديل (PFI)، التي واجهت صعوبات بسبب الشبكات التعاونية الكامنة لبيانات الجينات. تؤكد الدراسة على أهمية اختيار طرق XAI المناسبة لنمذجة التنبؤ النسخي، موصية باستخدام DeepSHAP لاكتشاف العلامات البيولوجية الجديدة وLRP للتحقق من الآليات السرطانية المعروفة. ومع ذلك، تشمل القيود التركيز الحصري على بيانات النسخ الجيني والحاجة إلى مزيد من الاستكشاف لدمج البيانات متعددة الأبعاد لتعزيز قابلية تعميم هذه النتائج. يجب أن تهدف الأبحاث المستقبلية إلى تحسين طرق XAI لتناسب تعقيدات مجموعات البيانات متعددة الأبعاد لتحسين قابليتها للتطبيق في الإعدادات السريرية.

Journal: Frontiers in Physiology, Volume: 17
DOI: https://doi.org/10.3389/fphys.2026.1830956
PMID: https://pubmed.ncbi.nlm.nih.gov/42099920
Publication Date: 2026-04-22
Author(s): 左熠熠 et al.
Primary Topic: Explainable Artificial Intelligence (XAI)

Overview

This section discusses the role of Explainable Artificial Intelligence (XAI) in addressing the challenges posed by the “black-box” nature of deep learning models in transcriptome-based cancer survival prediction. The authors highlight the absence of systematic benchmarking for XAI frameworks specifically designed for high-dimensional survival data. To address this, they evaluated six representative XAI methods—categorized as gradient-based, propagation-based, and perturbation-based—using a Self-Normalizing Neural Network (SNN) as the baseline model. The analysis utilized 6,248 samples from 15 cancer types sourced from The Cancer Genome Atlas (TCGA) within a unified framework.

The evaluation metrics focused on three dimensions: prognostic factor enrichment (measured by univariate Cox regression significance), biological consistency (validated against four authoritative databases, including OpenTargets), and explanation stability (assessed using the Kuncheva Index). The findings revealed that DeepSHAP outperformed the other methods, identifying the most statistically significant prognostic factors while ensuring high explanation stability. LRP (Layer-wise Relevance Propagation) demonstrated slightly lower prognostic specificity but excelled in aligning with biological databases, making it effective for validating biological relevance. Conversely, the perturbation-based method, PFI (Permutation Feature Importance), exhibited significant limitations, particularly in handling feature collinearity, resulting in low stability. The study concludes by establishing explanation stability as a reliable indicator of biological validity in XAI, thereby providing a framework for selecting effective AI explanation tools in precision medicine.

Methods

In this section, the authors detail their experimental design and statistical framework for evaluating the performance of explainable artificial intelligence (XAI) methods in identifying prognostic factors from pan-cancer transcriptomic data. They employed a 10× repeated 5-fold cross-validation strategy, resulting in 50 independent models per cancer cohort. The internal training and validation sets were used for monitoring early stopping, and 100 Bootstrap samplings were conducted to estimate confidence intervals for correlation analyses. The study found that DeepSHAP significantly outperformed other XAI methods, such as PFI and LRP, in identifying clinically relevant prognostic genes across 15 cancer types, with DeepSHAP yielding a median of 37.5 significant genes compared to reductions of 23.7% and 31.6% for PFI and LRP, respectively (Wilcoxon signed-rank test, FDR-adjusted p < 0.001). The analysis also revealed that gradient- and propagation-based methods demonstrated greater stability in feature attribution compared to perturbation-based approaches, as evidenced by higher Kuncheva Index scores (Friedman test, p < 0.001). DeepSHAP consistently ranked among the top methods across various cancer types, particularly in Glioblastoma and Lower Grade Glioma, where it identified the highest number of prognostic factors with a compact distribution across models. Furthermore, a strong positive correlation was observed between model predictive performance (C-index) and explanation stability, particularly in cohorts with robust predictive capabilities. This suggests that stable feature attributions are more likely to reflect biologically meaningful prognostic signals, while unstable methods like PFI may capture random variations rather than significant biological insights.

Results

The “Results” section of the research paper presents the findings derived from the implemented software and experimental environment. Key outcomes indicate the effectiveness of the proposed methodology in achieving the desired objectives. The results demonstrate significant improvements in performance metrics compared to baseline models, highlighting the robustness of the approach under various conditions.

Statistical analyses confirm the reliability of the results, with p-values indicating strong significance. Additionally, visual representations of the data, such as graphs and charts, illustrate the comparative advantages of the new method, reinforcing its applicability in real-world scenarios. Overall, the findings contribute valuable insights into the field and suggest avenues for future research.

Discussion

In this study, we analyzed a pan-cancer dataset comprising 6,248 patient samples across 15 cancer types, utilizing a deep survival prediction framework based on Scaled Exponential Linear Units (SELU) and Alpha Dropout to address the challenges of high-dimensional transcriptomic data. Our model achieved a median concordance index (C-index) of 0.622, demonstrating robust performance in survival prediction. We employed various explainable artificial intelligence (XAI) methods, including Integrated Gradients, DeepLIFT, and Layer-wise Relevance Propagation (LRP), to evaluate the interpretability of our model. Notably, LRP exhibited the highest biological consistency with established cancer gene databases, while DeepSHAP emerged as the most effective method for identifying prognostic biomarkers, balancing prognostic enrichment, biological consistency, and explanation stability.

Our findings indicate that propagation-based methods, particularly DeepSHAP, are superior in high-dimensional genomic contexts, as they avoid generating biologically implausible inputs, unlike perturbation-based methods such as Permutation Feature Importance (PFI), which struggled due to the inherent co-expression networks of gene data. The study underscores the importance of selecting appropriate XAI methods for transcriptomic prognostic modeling, recommending DeepSHAP for novel biomarker discovery and LRP for validating established cancer mechanisms. However, limitations include the exclusive focus on transcriptomic data and the need for further exploration of multi-omics integration to enhance the generalizability of these findings. Future research should aim to optimize XAI methods for the complexities of multi-omics datasets to improve their applicability in clinical settings.

شارك: