Optimal dose selection in phase I/II dose finding trial with contextual bandits: a case study and practical recommendations
Jixian Wang & Ram C. Tiwari
What the paper says
Dose selection is a key decision to make in the early phase of drug development. Classical phase I/II dose-finding trials randomly assign a few doses and select the best among them. Response-adaptive assignment designs are more efficient but are still far from optimal. Recently, some researchers used machine learning (ML) methods such as contextual bandits (CB) to find the "optimal" dose and to investigate the asymptotic properties of the methods. We present a case study for oncology phase I/II dose-finding trial designs using Thompson sampling and Bayesian bootstrap for CB with either modeling clinical utility directly or jointly modeling efficacy and safety. We focus on practical questions such as the number of interim analyses to conduct and whether we should model the utility directly, jointly model efficacy and safety which compose the utility, or use a model independent approach such as multi-armed bandits, but not for a specific compound or tumor type. We also consider how to use weak informative prior information. We conducted an extensive simulation study and compared different combinations of design settings and modeling methods, under several feasible scenarios of the dose-response relationship. Based on simulation results, we make practical recommendations for the use of the proposed ML approach for phase I/II dose-finding trial designs.
4 citations
Evidence weight
Balanced mode · F 0.40 / M 0.15 / V 0.05 / R 0.40
| F · citation impact | 0.37 × 0.4 = 0.15 |
| M · momentum | 0.60 × 0.15 = 0.09 |
| V · venue signal | 0.50 × 0.05 = 0.03 |
| R · text relevance † | 0.50 × 0.4 = 0.20 |
† Text relevance is estimated at 0.50 on the detail page — for your query’s actual relevance score, open this paper from a search result.