Offline Policy Evaluation and Learning with Harm Constraints

Offline policy learning aims to optimize individualized decisions using historical data. However, conventional methods primarily focus on maximizing expected rewards while neglecting individual-level counterfactual harm—cases where the assigned treatment leads to worse outcomes than the control. This can result in overly aggressive policies that harm some individuals. In this article, we propose a counterfactual harm-aware policy learning framework that effectively balances expected reward with individual-level harm. To quantify harm, we introduce two measures: treatment harm rate (THR), capturing the proportion of individuals harmed, and treatment harm quantity (THQ), capturing the severity of harm, particularly for continuous outcomes. A key challenge is that both THR and THQ depend on the joint distribution of potential outcomes, which is generally unidentifiable, even in randomized controlled trials. To address this, we introduce an additive latent variable model under which we establish the identifiability of THR and THQ and develop the corresponding estimators. These estimators are then used to learn policies under explicit harm constraints. Extensive experiments demonstrate the effectiveness of the proposed approach.

Authors

Institutions

Publication Details

Journal
Entropy
Published
2026-09-25
DOI
https://doi.org/10.3390/e28101058
Primary Topic
Advanced Causal Inference Techniques
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Offline Policy Evaluation and Learning with Harm Constraints

Qinwei Yang, Jile Chaoge, Jingyi Li, Peng Wu
Entropy
Advanced Causal Inference Techniques
article

Offline Policy Evaluation and Learning with Harm Constraints

Qinwei Yang, Jile Chaoge, Jingyi Li, Peng Wu
article en

Abstract

Offline policy learning aims to optimize individualized decisions using historical data. However, conventional methods primarily focus on maximizing expected rewards while neglecting individual-level counterfactual harm—cases where the assigned treatment leads to worse outcomes than the control. This can result in overly aggressive policies that harm some individuals. In this article, we propose a counterfactual harm-aware policy learning framework that effectively balances expected reward with individual-level harm. To quantify harm, we introduce two measures: treatment harm rate (THR), capturing the proportion of individuals harmed, and treatment harm quantity (THQ), capturing the severity of harm, particularly for continuous outcomes. A key challenge is that both THR and THQ depend on the joint distribution of potential outcomes, which is generally unidentifiable, even in randomized controlled trials. To address this, we introduce an additive latent variable model under which we establish the identifiability of THR and THQ and develop the corresponding estimators. These estimators are then used to learn policies under explicit harm constraints. Extensive experiments demonstrate the effectiveness of the proposed approach.

EntropyVol. 28(10)
University of International Business and Economics (CN), Beijing Technology and Business University (CN)
Peace, Justice and strong institutions
Openalex Percentile: Top 8%
Advanced Causal Inference Techniques
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Offline Policy Evaluation and Learning with Harm Constraints — Qinwei Yang, Jile Chaoge, et al. · Entropy (2026) | TGRS Research Map | TGRS