ARMA-Design: Optimal Treatment Allocation Strategies for A/B Testing in Partially Observable Experiments
Online experiments are frequently employed in technological companies to evaluate the performance of a newly developed policy, product, or treatment relative to a baseline control. In many applications, the experimental units receive a sequence of treatments over time. To handle these time-dependent settings, existing A/B testing solutions typically assume a fully observable experimental environment that satisfies the Markov condition. However, this assumption often does not hold in practice. This paper studies the optimal design for A/B testing in partially observable online experiments. We introduce a controlled (vector) autoregressive moving average model to capture partial observability. We introduce a small signal asymptotic framework to simplify the calculation of asymptotic mean squared errors of average treatment effect estimators under various designs. We develop two algorithms to estimate the optimal design: one utilizing constrained optimization and the other employing reinforcement learning. We demonstrate the superior performance of our designs using two dispatch simulators that realistically mimic the behaviors of drivers and passengers to create virtual environments, along with two real datasets from a ride-sharing company.
Authors
- Hongtu Zhu
- Ke Sun
- Linglong Kong
- Chengchun Shi
Institutions
- University of North Carolina at Chapel Hill (US)
- University of Alberta (CA)
- London School of Economics and Political Science (GB)
Publication Details
- Journal
- Journal of the American Statistical Association
- Published
- 2026-09-28
- DOI
- https://doi.org/10.1080/01621459.2026.2735068
- Primary Topic
- Advanced Causal Inference Techniques
- Type
- article
- Field-Weighted Citation Impact
- 0.00