BUDS: Benchmark Uncertainty Design Selection for two-stage single-arm phase II trials

Simon two-stage designs for binary endpoints and their time-to-event analogues, including the restricted-Kwak and Jung method, rely on a fixed historical benchmark. In practice, historical benchmarks are often uncertain due to small samples, population heterogeneity, changing eligibility criteria, and evolving standards of care. When the clinically relevant benchmark exceeds the value used to calibrate the design, Type I error rates can be substantially inflated, leading to costly advancement of ineffective treatments. Existing design-selection criteria generally optimize efficiency at a single benchmark without accounting for this uncertainty. We propose Benchmark Uncertainty Design Selection (BUDS), a framework for selecting two-stage designs for both binary and time-to-event (TTE) endpoints. Rather than relying on a single planning benchmark, BUDS addresses benchmark uncertainty by specifying a plausible range of response rates for binary endpoints or survival probabilities at a prespecified timepoint for TTE endpoints. Among feasible designs that satisfy both Type I error ( \(\:\alpha\:\) ) and Type II error ( \(\:\beta\:\) ) constraints, BUDS-Least-Regret minimizes the maximum difference in expected sample size from the benchmark-specific optimal design, whereas BUDS-Avg-EN minimizes average expected sample size across the range. By monotonicity, Type I error rate is controlled at the upper bound of the benchmark range. Across representative scenarios, the Type I error rate increases substantially when true benchmarks exceed their planning values using single-benchmark objectives, reaching 0.407 for Simon Optimal design (planning/true benchmark: 0.100 vs. 0.250; \(\:\alpha\:\) = 0.10, \(\:\beta\:\) = 0.10) and 0.165 for restricted-Kwak and Jung design (planning/true benchmark: 0.500 vs. 0.673; \(\:\alpha\:\) = 0.05, \(\:\beta\:\) = 0.20). In contrast, BUDS-selected designs maintain Type I error rate at or below the nominal level across the benchmark range, while the proposed objectives provide explicit criteria for prioritizing regret versus efficiency. BUDS provides a transparent framework for selecting designs when historical benchmarks are uncertain. It makes the trade-off between robustness and efficiency explicit while controlling Type I error across the plausible benchmark range. The framework applies to both binary and time-to-event endpoints in two-stage single-arm Phase II trial planning and is implemented in the open-source BUDS R package and interactive Shiny app.

Authors

Institutions

Publication Details

Journal
BMC Medical Research Methodology
Published
2026-09-24
DOI
https://doi.org/10.1186/s12874-026-02994-y
Primary Topic
Statistical Methods in Clinical Trials
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

BUDS: Benchmark Uncertainty Design Selection for two-stage single-arm phase II trials

Zhuoli Jin, Fei Ye, Rebecca Irlmeier
BMC Medical Research Methodology
Statistical Methods in Clinical Trials
article

BUDS: Benchmark Uncertainty Design Selection for two-stage single-arm phase II trials

Zhuoli Jin, Fei Ye, Rebecca Irlmeier
article en

Abstract

Simon two-stage designs for binary endpoints and their time-to-event analogues, including the restricted-Kwak and Jung method, rely on a fixed historical benchmark. In practice, historical benchmarks are often uncertain due to small samples, population heterogeneity, changing eligibility criteria, and evolving standards of care. When the clinically relevant benchmark exceeds the value used to calibrate the design, Type I error rates can be substantially inflated, leading to costly advancement of ineffective treatments. Existing design-selection criteria generally optimize efficiency at a single benchmark without accounting for this uncertainty. We propose Benchmark Uncertainty Design Selection (BUDS), a framework for selecting two-stage designs for both binary and time-to-event (TTE) endpoints. Rather than relying on a single planning benchmark, BUDS addresses benchmark uncertainty by specifying a plausible range of response rates for binary endpoints or survival probabilities at a prespecified timepoint for TTE endpoints. Among feasible designs that satisfy both Type I error ( \(\:\alpha\:\) ) and Type II error ( \(\:\beta\:\) ) constraints, BUDS-Least-Regret minimizes the maximum difference in expected sample size from the benchmark-specific optimal design, whereas BUDS-Avg-EN minimizes average expected sample size across the range. By monotonicity, Type I error rate is controlled at the upper bound of the benchmark range. Across representative scenarios, the Type I error rate increases substantially when true benchmarks exceed their planning values using single-benchmark objectives, reaching 0.407 for Simon Optimal design (planning/true benchmark: 0.100 vs. 0.250; \(\:\alpha\:\) = 0.10, \(\:\beta\:\) = 0.10) and 0.165 for restricted-Kwak and Jung design (planning/true benchmark: 0.500 vs. 0.673; \(\:\alpha\:\) = 0.05, \(\:\beta\:\) = 0.20). In contrast, BUDS-selected designs maintain Type I error rate at or below the nominal level across the benchmark range, while the proposed objectives provide explicit criteria for prioritizing regret versus efficiency. BUDS provides a transparent framework for selecting designs when historical benchmarks are uncertain. It makes the trade-off between robustness and efficiency explicit while controlling Type I error across the plausible benchmark range. The framework applies to both binary and time-to-event endpoints in two-stage single-arm Phase II trial planning and is implemented in the open-source BUDS R package and interactive Shiny app.

BMC Medical Research Methodology
University of Miami (US), Sylvester Comprehensive Cancer Center (US)
Openalex Percentile: Top 8%
Statistical Methods in Clinical Trials
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.