DOI: https://doi.org/10.1093/jrsssc/qlag013
تاريخ النشر: 2026-02-20
المؤلف: Zhenyun Du
الموضوع الرئيسي: تقنيات الاستدلال السببي المتقدمة
نظرة عامة
إطار عمل Buckley-James (BJ) Boost Q-learning المقترح يعالج تحدي تقدير أنظمة العلاج الديناميكية المثلى من نتائج البقاء المقيدة من اليمين في التجارب السريرية العشوائية الطولية. هذه الطريقة ذات صلة خاصة لدعم قرارات العلاج المحددة للمرضى في السيناريوهات التي قد تكون فيها بيانات المتابعة غير مكتملة وتأثيرات المتغيرات المحتملة غير خطية. من خلال دمج نمذجة وقت الفشل المعجل مع تقنيات التعزيز التكرارية – باستخدام متعلمين أساسيين مرنين مثل المربعات الصغرى المكونة وأشجار الانحدار – يعمل BJ Boost Q-learning ضمن إطار عمل Q-learning المضاد للوقائع. هذه الطريقة تقوم بنمذجة وقت البقاء الشرطي مباشرة، متجاوزة قيود افتراض المخاطر النسبية وتسهيل تقدير وظائف Q الخاصة بالمراحل وقواعد القرار الفردية تحت افتراضات النتائج المحتملة القياسية.
بالمقارنة مع Q-learning التقليدي القائم على نموذج كوكس، الذي يكون حساسًا للمخاطر غير النسبية وسوء تحديد النموذج، يقدم BJ Boost Q-learning بديلاً أكثر قوة ومرونة لتعلم الأنظمة. تظهر دراسات المحاكاة، جنبًا إلى جنب مع تحليلات تجربة ACTG175 لفيروس نقص المناعة البشرية وتجربة CALGB 8923 لسرطان الدم ذات المرحلتين، أن هذا الإطار يعزز دقة قرارات العلاج ويحقق تباينات مضادة للوقائع أكثر استقرارًا ضمن المشاركين. هذا مفيد بشكل خاص في الإعدادات متعددة المراحل، حيث يمكن أن يؤثر تراكم خطأ التقدير والتحيز بشكل كبير على النتائج.
مقدمة
في مجال الرعاية الصحية المعاصرة، ظهرت استراتيجيات العلاج الفردية كنهج محوري لتحسين نتائج المرضى، خاصة في سياق البيانات المقيدة المعقدة حيث قد تكون أوقات الأحداث غير مشهودة جزئيًا. تواجه طرق تحليل البقاء التقليدية، مثل انحدار كوكس، تحديات بسبب افتراضات المخاطر النسبية وصعوبة تفسير نسب المخاطر. يقدم نموذج وقت الفشل المعجل (AFT) بديلاً قويًا، مما يسمح بإطار عمل أكثر قابلية للتفسير من خلال نمذجة وقت البقاء مباشرة. طريقة Buckley-James (BJ)، وهي مقدر شبه بارامتري لنموذج AFT، تستوعب آليات التقييد التعسفية وقد تم توسيعها من خلال طرق معاقبة للتعامل مع البيانات عالية الأبعاد.
لزيادة تعزيز إطار عمل BJ، تم تطوير خوارزمية BJ Boosting، التي تقوم بتحسين الوظائف التنبؤية بشكل تكراري باستخدام متعلمين أساسيين مرنين. تعالج هذه الطريقة بشكل فعال التقييد من اليمين وتلتقط العلاقات المعقدة وغير الخطية، مما يظهر أداءً متفوقًا في الكفاءة الحسابية ودقة التقدير مقارنة بالطرق الكلاسيكية. بالإضافة إلى ذلك، يهدف دمج Q-learning ضمن إطار عمل مضاد للوقائع إلى تحسين استراتيجيات العلاج الشخصية من خلال التكيف مع الحالات الصحية المتطورة للمرضى. هذه الطريقة الجديدة BJ Boost Q-learning ذات صلة خاصة في التجارب السريرية، حيث تسهل تقدير النتائج المحتملة تحت استراتيجيات علاجية متنوعة، مما يحسن تقييمات أنظمة العلاج الديناميكية ويعزز اتخاذ القرارات السريرية المخصصة لاحتياجات المرضى الفردية.
نقاش
في هذا القسم، يناقش المؤلفون إطار تجربة سريرية عشوائية متسلسلة مصممة لتقدير قرارات العلاج المثلى من خلال نهج Q-learning متعدد المراحل، مع التركيز بشكل خاص على بيانات البقاء المقيدة من اليمين. تتضمن المنهجية تعزيز Buckley-James لتقدير النتائج المقيدة وبناء نتائج زائفة، والتي تعتبر ضرورية لتقدير وظائف Q. يحدد المؤلفون المتغيرات والعلاقات الرئيسية، بما في ذلك تاريخ العلاج، والنتائج المحتملة، ومؤشرات التقييد، بينما يضعون افتراضات التعرف اللازمة لصحة تقديراتهم.
تقيم دراسة المحاكاة أداء طرق تقدير قيمة Q المختلفة، بما في ذلك الانحدار الخطي التقليدي وتقنيات التعزيز المتقدمة، عبر إعدادات ذات مرحلة واحدة وذوات مرحلتين. تشير النتائج إلى أن طرق التعزيز، وخاصة نهج BJ-Tree، تتفوق بشكل كبير على النماذج القائمة على كوكس من حيث دقة قرارات العلاج، حيث تحقق دقة متوسطة تتجاوز 91% في الإعدادات ذات المرحلة الواحدة وأكثر من 84% في الإعدادات ذات المرحلتين. تؤكد النتائج على أهمية استخدام نماذج مرنة وغير بارامترية لالتقاط تفاعلات المتغيرات العلاجية المعقدة بشكل فعال في التجارب السريرية، خاصة في ظل ظروف التقييد العالي. يدعم تطبيق هذه المنهجية على مجموعة بيانات ACTG175 فعالية العلاج المركب مقارنة بالعلاج الأحادي، مما يتماشى مع النتائج السريرية السابقة.
DOI: https://doi.org/10.1093/jrsssc/qlag013
Publication Date: 2026-02-20
Author(s): Zhenyun Du
Primary Topic: Advanced Causal Inference Techniques
Overview
The proposed Buckley-James (BJ) Boost Q-learning framework addresses the challenge of estimating optimal dynamic treatment regimes from right-censored survival outcomes in longitudinal randomized clinical trials. This method is particularly relevant for supporting patient-specific treatment decisions in scenarios where follow-up data may be incomplete and covariate effects are potentially nonlinear. By integrating accelerated failure time modeling with iterative boosting techniques—utilizing flexible base learners such as componentwise least squares and regression trees—BJ Boost Q-learning operates within a counterfactual Q-learning framework. This approach directly models conditional survival time, circumventing the limitations of the proportional hazards assumption and facilitating the estimation of stage-specific Q-functions and individualized decision rules under standard potential outcomes assumptions.
In contrast to traditional Cox-based Q-learning, which is sensitive to nonproportional hazards and model misspecification, BJ Boost Q-learning offers a more robust and flexible alternative for regime learning. Simulation studies, along with analyses of the ACTG175 HIV trial and the CALGB 8923 two-stage leukemia trial, demonstrate that this framework enhances treatment decision accuracy and yields more stable within-participant counterfactual contrasts. This is particularly beneficial in multistage settings, where the compounding of estimation error and bias can significantly impact outcomes.
Introduction
In the realm of contemporary healthcare, individualized treatment strategies have emerged as a pivotal approach to optimize patient outcomes, particularly in the context of complex censored data where event times may be partially unobserved. Traditional survival analysis methods, such as Cox regression, face challenges due to the assumptions of proportional hazards and the difficulty in interpreting hazard ratios. The accelerated failure time (AFT) model presents a robust alternative, allowing for a more interpretable framework by directly modeling survival time. The Buckley-James (BJ) method, a semiparametric estimator for the AFT model, accommodates arbitrary censoring mechanisms and has been extended through penalized methods to handle high-dimensional data.
To further enhance the BJ framework, the BJ Boosting algorithm has been developed, which iteratively refines predictive functions using flexible base learners. This approach effectively addresses right-censoring and captures complex, non-linear relationships, demonstrating superior performance in computational efficiency and estimation accuracy compared to classical methods. Additionally, the integration of Q-learning within a counterfactual framework aims to optimize personalized treatment strategies by adapting to the evolving health states of patients. This novel BJ Boost Q-learning approach is particularly relevant in clinical trials, where it facilitates the estimation of potential outcomes under various treatment strategies, thereby improving dynamic treatment regime evaluations and enhancing clinical decision-making tailored to individual patient needs.
Discussion
In this section, the authors discuss a sequentially randomized clinical trial framework designed to estimate optimal treatment decisions through a multi-stage Q-learning approach, particularly addressing right-censored survival data. The methodology incorporates Buckley-James boosting for imputing censored outcomes and constructing pseudo outcomes, which are essential for estimating Q-functions. The authors define key variables and relationships, including treatment histories, potential outcomes, and censoring indicators, while establishing identifiability assumptions necessary for the validity of their estimations.
The simulation study evaluates the performance of various Q-value estimation methods, including traditional linear regression and advanced boosting techniques, across single-stage and two-stage settings. Results indicate that the boosting methods, particularly the BJ-Tree approach, significantly outperform Cox-based models in terms of treatment decision accuracy, achieving median accuracies above 91% in single-stage settings and over 84% in two-stage settings. The findings emphasize the importance of employing flexible, nonparametric models to effectively capture complex covariate-treatment interactions in clinical trials, particularly under conditions of high censoring. The application of this methodology to the ACTG175 dataset further supports the efficacy of combination therapy over monotherapy, aligning with previous clinical findings.
