Counterfactual survival Q-learning via Buckley–James boosting with applications to ACTG 175 and CALGB 8923

Jeongjin Lee & Jong-Min Kim

Journal of the Royal Statistical Society. Series C: Applied Statistics2026https://doi.org/10.1093/jrsssc/qlag013article
AJG 3
Weight
0.37

Abstract

We propose a Buckley–James (BJ) Boost Q-learning framework for estimating optimal dynamic treatment regimes from right censored survival outcomes in longitudinal randomized clinical trials, motivated by the clinical need to support patient specific treatment decisions when follow up is incomplete and covariate effects may be nonlinear. The method combines accelerated failure time modelling with iterative boosting using flexible base learners, including componentwise least squares and regression trees, within a counterfactual Q-learning framework. By modelling conditional survival time directly, BJ Boost Q-learning avoids the proportional hazards assumption, yields clinically interpretable time scale contrasts, and enables estimation of stage specific Q-functions and individualized decision rules under standard potential outcomes assumptions. In contrast to Cox-based Q-learning, which relies on hazard modelling and can be sensitive to nonproportional hazards and model misspecification, our approach provides a robust and flexible alternative for regime learning. Simulation studies and analyses of the ACTG175 HIV trial and the CALGB 8923 two-stage leukaemia trial show that BJ Boost Q-learning improves treatment decision accuracy and produces more stable within participant counterfactual contrasts, particularly in multistage settings where estimation error and bias can compound across stages.

1 citation

Open via your library →

Cite this paper

https://doi.org/https://doi.org/10.1093/jrsssc/qlag013

Or copy a formatted citation

@article{jeongjin2026,
  title        = {{Counterfactual survival Q-learning via Buckley–James boosting with applications to ACTG 175 and CALGB 8923}},
  author       = {Jeongjin Lee & Jong-Min Kim},
  journal      = {Journal of the Royal Statistical Society. Series C: Applied Statistics},
  year         = {2026},
  doi          = {https://doi.org/https://doi.org/10.1093/jrsssc/qlag013},
}

Paste directly into BibTeX, Zotero, or your reference manager.

Flag this paper

Counterfactual survival Q-learning via Buckley–James boosting with applications to ACTG 175 and CALGB 8923

Flags are reviewed by the Arbiter methodology team within 5 business days.


Evidence weight

0.37

Balanced mode · F 0.40 / M 0.15 / V 0.05 / R 0.40

F · citation impact0.16 × 0.4 = 0.06
M · momentum0.53 × 0.15 = 0.08
V · venue signal0.50 × 0.05 = 0.03
R · text relevance †0.50 × 0.4 = 0.20

† Text relevance is estimated at 0.50 on the detail page — for your query’s actual relevance score, open this paper from a search result.