Offline Policy Evaluation and Learning with Harm Constraints
Offline policy learning aims to optimize individualized decisions using historical data. However, conventional methods primarily focus on maximizing expected rewards while neglecting individual-level counterfactual harm—cases where the assigned treatment leads to worse outcomes than the control. This can result in overly aggressive policies that harm some individuals. In this article, we propose a counterfactual harm-aware policy learning framework that effectively balances expected reward with individual-level harm. To quantify harm, we introduce two measures: treatment harm rate (THR), capturing the proportion of individuals harmed, and treatment harm quantity (THQ), capturing the severity of harm, particularly for continuous outcomes. A key challenge is that both THR and THQ depend on the joint distribution of potential outcomes, which is generally unidentifiable, even in randomized controlled trials. To address this, we introduce an additive latent variable model under which we establish the identifiability of THR and THQ and develop the corresponding estimators. These estimators are then used to learn policies under explicit harm constraints. Extensive experiments demonstrate the effectiveness of the proposed approach.
Authors
- Qinwei Yang (ORCID: https://orcid.org/0009-0004-4754-8994)
- Jile Chaoge (ORCID: https://orcid.org/0009-0009-1245-8602)
- Jingyi Li
- Peng Wu
Institutions
- University of International Business and Economics (CN)
- Beijing Technology and Business University (CN)
Publication Details
- Journal
- Entropy
- Published
- 2026-09-25
- DOI
- https://doi.org/10.3390/e28101058
- Primary Topic
- Advanced Causal Inference Techniques
- Type
- article
- Field-Weighted Citation Impact
- 0.00