دمج الأشجار الاحتمالية والشبكات السببية للبيانات السريرية والوبائية
Integrating probabilistic trees and causal networks for clinical and epidemiological data

شارك:
المجلة: Artificial Intelligence in Medicine، المجلد: 173
DOI: https://doi.org/10.1016/j.artmed.2026.103350
PMID: https://pubmed.ncbi.nlm.nih.gov/41512598
تاريخ النشر: 2026-01-06
المؤلف: Sheresh Zahoor وآخرون
الموضوع الرئيسي: النمذجة البايزية والاستدلال السببي

نظرة عامة

تقدم البحث إطار عمل الدمج السببي الاحتمالي (PCF)، الذي يدمج الشبكات البايزية السببية (CBNs) وأشجار الاحتمالات (PTrees) لتعزيز اتخاذ القرارات في الرعاية الصحية من خلال التنبؤ بالنتائج واستكشاف آثار التدخلات. يلتقط إطار عمل PCF العلاقات السببية، مما يسمح بتحديد تأثيرات العوامل ومحاكاة السيناريوهات الافتراضية. تم تقييمه على مجموعات بيانات متنوعة بما في ذلك MIMIC-IV، دراسة قلب فرامينغهام، وBRFSS، حيث يظهر PCF أداءً تنبؤياً قابلاً للمقارنة مع نماذج التعلم الآلي التقليدية مع توفير قابلية تفسير محسّنة من خلال تحليل الحساسية وSHapley Additive exPlanations (SHAP). تساعد هذه القابلية التفسيرية ذات الطبقتين الأطباء في فهم كل من المسارات السببية على المستوى الكلي والتنبؤات الفردية على المستوى الدقيق.

يؤكد البحث على أن قوة PCF تكمن في قدرته على تسهيل التحليل المضاد للحقائق، مما يمكّن الأطباء من استكشاف التدخلات المحتملة وآثارها، وبالتالي سد الفجوة بين الحدس السريري والرؤى المستندة إلى البيانات. بينما يحقق دقة تنبؤية مماثلة للنماذج المعتمدة، يولي PCF الأولوية لقابلية التفسير السببي، مما قد يؤدي إلى تقليل طفيف في الأداء التنبؤي. تسمح مرونة الإطار بالحصول على رؤى على كل من المستويات الفردية والسكانية، مما يدعم اتخاذ قرارات مخصصة. ومع ذلك، يعترف المؤلفون بالقيود مثل الكفاءة الحاسوبية، والافتراضات المتعلقة بالقدرة على التعرف السببي، والتحديات في مشاركة البيانات. يتم تشجيع الأبحاث المستقبلية لمعالجة هذه القضايا والتحقق من صحة الإطار عبر سياقات الرعاية الصحية الأوسع، بهدف تعزيز اتخاذ القرارات السريرية المستندة إلى الأدلة وتحسين نتائج رعاية المرضى.

مقدمة

تستعرض مقدمة البحث قيود نماذج التعلم الآلي التقليدية في الرعاية الصحية، ولا سيما عدم قدرتها على معالجة الأسئلة السببية المتعلقة بنتائج المرضى وآثار العلاج. لمعالجة هذه الفجوة، يقترح المؤلفون إطار عمل الدمج السببي الاحتمالي (PCF)، الذي يدمج الشبكات البايزية السببية (CBNs) ومجموعات أشجار الاحتمالات (PTrees). توفر CBNs نهجًا منظمًا للاستدلال السببي من خلال الرسوم البيانية الموجهة غير الدائرية (DAGs)، بينما تعزز PTrees القابلية للتفسير وتتوافق مع سير العمل السريري من خلال تقديم قواعد قرار واضحة ومحددة للسياق. يهدف إطار عمل PCF إلى تحسين التفكير السببي ومحاكاة التدخلات من خلال الاستفادة من الهيكل السببي المستفاد من CBNs لإبلاغ ترتيب المتغيرات في PTrees، مما يقلل من الاعتماد على المدخلات الخبيرة الذاتية.

يؤكد المؤلفون على أهمية هذا النهج الهجين في الإعدادات السريرية، حيث تكون البيانات الملاحظة شائعة، وغالبًا ما تكون الدراسات التجريبية غير عملية. من خلال تسهيل الاستدلال السببي مباشرة من البيانات الملاحظة، يدعم PCF الأطباء في فهم ليس فقط النتائج المحتملة ولكن أيضًا الآليات وراءها. تم تصميم الإطار ليكون قابلاً للتكيف، مما يسمح بإجراء تعديلات مستنيرة من قبل الأطباء على ترتيب المتغيرات مع الحفاظ على سلامة الافتراضات السببية. تعزز هذه القابلية للتكيف من قابلية تفسير النموذج وفعاليته التشغيلية، مما يعزز في النهاية بيئة تعاونية لنقل المعرفة بين مؤسسات الرعاية الصحية. تختتم الورقة بتحديد تطبيقات الإطار في التنبؤ بمدة الإقامة في وحدة العناية المركزة، وتقييم مخاطر أمراض القلب المزمنة، وتحليل بداية مرض السكري، مما يبرز مساهماته في تعزيز التفكير السببي في تحليلات الرعاية الصحية.

الطرق

تستعرض هذه القسم منهجية PCF (إطار العمل السببي الاحتمالي)، الذي يدمج الشبكات البايزية السببية (CBNs) مع مجموعات أشجار الاحتمالات (PTrees) لتعزيز الاستدلال السببي في اتخاذ القرارات السريرية. يبدأ بمقارنة PCF بأساليب الاستدلال السببي الحالية – مقدرات غير مرتبطة بالنموذج، نماذج التعلم العميق، ونماذج السببية الهيكلية – مع تسليط الضوء على نقاط القوة والقيود لكل منها. بينما توفر المقدرات غير المعلمية مثل الغابات السببية وأشجار الانحدار البايزية نماذج فعالة لتباين العلاج، إلا أنها تفتقر إلى التفكير المضاد للحقائق بشكل صريح. تقدم طرق التعلم العميق للتمثيل السببي، مثل TARNet وCounterfactual Variational Autoencoders، مزايا في البيانات عالية الأبعاد ولكنها غالبًا ما تعمل كصناديق سوداء مع قابلية تفسير محدودة. توفر النماذج السببية الهيكلية استدلالًا سببيًا صارمًا ولكنها تتطلب افتراضات قوية قد تكون صعبة التحقق.

يعالج إطار عمل PCF هذه القيود من خلال اعتماد بناء شجري قائم على الوحدات يعزز الكفاءة الحاسوبية وقابلية التفسير دون الحاجة إلى معادلات هيكلية محددة مسبقًا. يجمع بين التوجيه الهيكلي لـ CBNs مع بنية احتمالية تحافظ على الاعتماد الشرطي، مما يسهل حساب التوزيعات التدخلية، \( P(Y | do(X)) \). يسمح الإطار بالتفكير المضاد للحقائق على مستوى المريض من خلال بناء شجري مزدوج، مما يدعم تحليلات “ماذا لو” الشفافة. بالإضافة إلى ذلك، يولد PCF ترتيب متغيرات مدفوع بالبيانات بناءً على الاعتماد الهيكلي، مما يحسن القابلية للتفسير وإمكانية التكرار مع السماح للأطباء بتنقيح الترتيب حسب الحاجة. من خلال المتوسط عبر عدة رسوم بيانية مرشحة، يعزز PCF من القوة والأداء التنبؤي، مما يوفر في النهاية دعامة أكثر استقرارًا وكفاءة لتطوير نماذج سببية ذات مغزى سريري.

النتائج

في قسم النتائج، يتم بدء عملية التقييم من خلال تحليل الحساسية في القسم 5.1، الذي يفحص تباين النتائج. يركز القسم 5.2 على الأداء التنبؤي لنموذج PCF، مقارنًا إياه بمختلف نماذج المعايير، سواء كانت قابلة للتفسير أو غير قابلة للتفسير، عبر ثلاث مجموعات بيانات. يتضمن هذا القسم أيضًا SHAP (SHapley Additive exPlanations) لتقييم قابلية تفسير النموذج.

بعد ذلك، يتناول القسم 5.3 التحليل التدخلي لتقييم آثار التدخلات المحتملة على نتائج النموذج. أخيرًا، يقدم القسم 5.4 تحليلًا مضادًا للحقائق، مما يوفر رؤى إضافية حول سلوك النموذج تحت سيناريوهات افتراضية مختلفة. تعزز هذه التحليلات مجتمعة من فهم فعالية PCF وقابلية تفسيره في سياقات متنوعة.

المناقشة

في قسم المناقشة، يقارن البحث بين التعلم الآلي التقليدي (ML) والتعلم الآلي السببي، مؤكدًا أنه بينما يتفوق التعلم الآلي التقليدي في المهام التنبؤية ضمن الرعاية الصحية، إلا أنه يقصر في معالجة الاستفسارات المضادة للحقائق وتقدير آثار العلاج. من ناحية أخرى، يسمح التعلم الآلي السببي باستكشاف سيناريوهات “ماذا لو”، مما يمكّن مقدمي الرعاية الصحية من اتخاذ قرارات مستنيرة بناءً على التدخلات المحتملة. يبرز المؤلفون أهمية الشبكات البايزية السببية (CBNs) وأشجار الاحتمالات (PTrees) كأدوات لنمذجة العلاقات المعقدة في بيانات الرعاية الصحية، مشيرين إلى أن CBNs تسهل فهم الاعتماد السببي من خلال الرسوم البيانية الموجهة غير الدائرية (DAGs)، بينما تقدم PTrees تمثيلًا أكثر بديهية للعلاقات الاحتمالية.

تناقش الورقة أيضًا التقدمات الأخيرة في التعلم الآلي السببي التي تعزز اتخاذ القرارات في الرعاية الصحية من خلال تقدير آثار العلاج ومحاكاة السيناريوهات المضادة للحقائق. تقدم إطار عمل الدمج السببي الاحتمالي (PCF)، الذي يدمج CBNs وPTrees لتحسين القوة وقابلية التفسير في النمذجة السببية. من خلال استخدام متوسط النماذج وتحليل الحساسية، يهدف إطار عمل PCF إلى تحسين الهياكل السببية مع استيعاب خبرة المجال. بالإضافة إلى ذلك، يعزز دمج SHapley additive explanations (SHAP) من شفافية النموذج من خلال تحديد أهمية الميزات، مما يعزز الثقة في تنبؤات النموذج. بشكل عام، يدعو المؤلفون إلى التطبيق العملي لتقنيات التعلم الآلي السببي في الإعدادات السريرية لسد الفجوة بين التحليلات التنبؤية والرؤى القابلة للتنفيذ في الرعاية الصحية.

Journal: Artificial Intelligence in Medicine, Volume: 173
DOI: https://doi.org/10.1016/j.artmed.2026.103350
PMID: https://pubmed.ncbi.nlm.nih.gov/41512598
Publication Date: 2026-01-06
Author(s): Sheresh Zahoor et al.
Primary Topic: Bayesian Modeling and Causal Inference

Overview

The research introduces the Probabilistic Causal Fusion (PCF) framework, which merges Causal Bayesian Networks (CBNs) and Probability Trees (PTrees) to enhance healthcare decision-making by not only predicting outcomes but also exploring the effects of interventions. The PCF framework captures causal relationships, allowing for the quantification of factor impacts and the simulation of hypothetical scenarios. Evaluated on diverse datasets including MIMIC-IV, Framingham Heart Study, and BRFSS, PCF demonstrates predictive performance comparable to traditional machine learning models while providing improved interpretability through sensitivity analysis and SHapley Additive exPlanations (SHAP). This dual-layered interpretability aids clinicians in understanding both macro-level causal pathways and micro-level individual predictions.

The study emphasizes that PCF’s strength lies in its ability to facilitate counterfactual analysis, enabling clinicians to explore potential interventions and their impacts, thus bridging the gap between clinical intuition and data-driven insights. While achieving similar predictive accuracy to established models, PCF prioritizes causal interpretability, which may result in slight reductions in predictive performance. The framework’s versatility allows for insights at both individual and population levels, supporting tailored decision-making. However, the authors acknowledge limitations such as computational efficiency, assumptions regarding causal identifiability, and challenges in data sharing. Future research is encouraged to address these issues and validate the framework across broader healthcare contexts, ultimately aiming to enhance evidence-based clinical decision-making and improve patient care outcomes.

Introduction

The introduction of the research paper outlines the limitations of traditional machine learning models in healthcare, particularly their inability to address causal questions regarding patient outcomes and treatment effects. To address this gap, the authors propose the Probabilistic Causal Fusion (PCF) framework, which integrates Causal Bayesian Networks (CBNs) and ensembles of Probability Trees (PTrees). CBNs provide a structured approach to causal inference through directed acyclic graphs (DAGs), while PTrees enhance interpretability and align with clinical workflows by offering clear, context-specific decision rules. The PCF framework aims to improve causal reasoning and intervention simulation by leveraging the causal structure learned from CBNs to inform variable ordering in PTrees, thereby reducing reliance on subjective expert input.

The authors emphasize the importance of this hybrid approach in clinical settings, where observational data is prevalent, and experimental studies are often impractical. By facilitating causal inference directly from observational data, PCF supports clinicians in understanding not only potential outcomes but also the mechanisms behind them. The framework is designed to be adaptable, allowing for clinician-informed adjustments to variable order while maintaining the integrity of causal assumptions. This adaptability enhances the model’s interpretability and operational efficacy, ultimately fostering a collaborative environment for knowledge transfer among healthcare institutions. The paper concludes by outlining the framework’s applications in predicting ICU length of stay, assessing chronic heart disease risk, and analyzing diabetes onset, highlighting its contributions to advancing causal reasoning in healthcare analytics.

Methods

The section outlines the PCF (Probabilistic Causal Framework) methodology, which integrates Causal Bayesian Networks (CBNs) with ensembles of Probabilistic Trees (PTrees) to enhance causal inference in clinical decision-making. It begins by comparing PCF to existing causal inference methods—model-agnostic estimators, deep representation learning models, and structural causal models—highlighting the strengths and limitations of each. While non-parametric estimators like Causal Forests and Bayesian Additive Regression Trees effectively model treatment heterogeneity, they lack explicit counterfactual reasoning. Deep causal representation learning methods, such as TARNet and Counterfactual Variational Autoencoders, offer advantages in high-dimensional data but often function as black boxes with limited interpretability. Structural Causal Models provide rigorous causal reasoning but require strong assumptions that can be challenging to validate.

The PCF framework addresses these limitations by employing a modular, tree-based construction that enhances computational efficiency and interpretability without necessitating predefined structural equations. It combines the structural guidance of CBNs with a probabilistic architecture that maintains conditional dependencies, facilitating the computation of interventional distributions, \( P(Y | do(X)) \). The framework allows for patient-level counterfactual reasoning through a twin-tree construction, thereby supporting transparent “what if” analyses. Additionally, PCF generates a data-driven variable ordering based on structural dependencies, improving interpretability and reproducibility while allowing clinicians to refine the ordering as needed. By averaging across multiple candidate graphs, PCF enhances robustness and predictive performance, ultimately providing a more stable and efficient scaffold for developing clinically meaningful causal models.

Results

In the Results section, the evaluation process is initiated with a sensitivity analysis in Section 5.1, which examines the variability of outcomes. Section 5.2 focuses on the predictive performance of the PCF model, comparing it to various benchmark models, both interpretable and non-interpretable, across three datasets. This section also incorporates SHAP (SHapley Additive exPlanations) to assess model interpretability.

Subsequently, Section 5.3 delves into interventional analysis to evaluate the impacts of potential interventions on the model’s outcomes. Finally, Section 5.4 presents a counterfactual analysis, offering additional insights into the model’s behavior under different hypothetical scenarios. These analyses collectively enhance the understanding of the PCF model’s efficacy and interpretability in various contexts.

Discussion

In the discussion section, the paper contrasts traditional machine learning (ML) with causal ML, emphasizing that while traditional ML excels in predictive tasks within healthcare, it falls short in addressing counterfactual inquiries and estimating treatment effects. Causal ML, on the other hand, allows for the exploration of “what if” scenarios, enabling healthcare providers to make informed decisions based on potential interventions. The authors highlight the importance of causal Bayesian networks (CBNs) and probability trees (PTrees) as tools for modeling complex relationships in healthcare data, noting that CBNs facilitate the understanding of causal dependencies through directed acyclic graphs (DAGs), while PTrees offer a more intuitive representation of probabilistic relationships.

The paper further discusses recent advancements in causal ML that enhance healthcare decision-making by estimating treatment effects and simulating counterfactual scenarios. It introduces the Probabilistic Causal Fusion (PCF) framework, which integrates CBNs and PTrees to improve robustness and interpretability in causal modeling. By employing model averaging and sensitivity analysis, the PCF framework aims to refine causal structures while accommodating domain expertise. Additionally, the integration of SHapley additive explanations (SHAP) enhances model transparency by quantifying feature importance, thereby fostering trust in the model’s predictions. Overall, the authors advocate for the practical application of causal ML techniques in clinical settings to bridge the gap between predictive analytics and actionable healthcare insights.

شارك: