A Decision-Theoretic Extension of Causal Unit Selection: Identifiability, Threshold Rules, and the Fallacy of Naive Benefit Rates.

ABSTRACT AND OVERVIEW This repository contains the official preprint and Python replication code for the research paper “A Decision-Theoretic Extension of Causal Unit Selection: Identifiability, Threshold Rules, and the Fallacy of Naive Benefit Rates” (Eskelinen, Version 4). The work formalizes and advances the principles of individualized causal decision-making by extending the foundational Li-Pearl (2021) benefit function into a decision-theoretic framework. While standard unit-selection frameworks focus on bounding unobserved counterfactual response types (Benefited, Always-taker, Never-taker, and Harm), this paper derives explicit operational decision rules for group selection under counterfactual uncertainty and finite-data constraints. Version 3 extended these results into an LLM Agent Autonomy Guardrail, deriving a mathematical gate for deciding when an AI agent may act autonomously and when it must defer to a human. Version 4 extends the Version 3 autonomy framework to strategic environments. The central question is what happens after the autonomy rule itself is deployed and becomes observable to the environment. Once agents, applicants, adversaries, or other systems can respond to the rule, the observed success margins used by the autonomy gate may become endogenous to the policy itself. This creates a Lucas-critique / Goodhart-type problem: parameters estimated before deployment need not remain invariant after actors begin optimizing against the deployed decision rule. Version 4 does not replace or reinvent the mathematics of T5-T9. Instead, it develops a strategic extension of the existing causal unit-selection and abstention structure and studies when re-estimation restores the safety properties of the original autonomy gate. CORE MATHEMATICAL AND THEORETICAL CONTRIBUTIONS Theorem 5 - Identifiability The target benefit function f(c) is point-identifiable purely from observable causal effects if and only if the structural master coefficient sigma = beta - gamma - theta + delta equals zero. When sigma is non-zero, counterfactual shifts can alter the true benefit function while leaving observable causal margins unchanged. Theorem 6 - Decision Criterion and the Abstain Rule Under fundamental counterfactual uncertainty where the Probability of Necessity and Sufficiency (PNS) is unknown, target groups c0 and c1 are strictly decidable if and only if |W1 - W0| > |sigma|. When this condition fails, the structural intervals overlap and the rule abstains from selection. Theorem 7 - Reliability Under Estimation For a true gap margin D = |W1 - W0| - |sigma| > 0, the decision rule yields zero active errors asymptotically, with the required empirical sample size scaling proportionally to 1 / D^2. When D <= 0, additional data alone cannot resolve the underlying structural uncertainty. Theorem 8 - Autonomy Criterion Extending T5-T7 to agent environments, an agent’s action is strictly and decidably better than deferring to a human baseline if and only if |gamma - delta| * |p_a - p_h| > |sigma|, where p_a denotes the agent success rate and p_h denotes the human success rate. The comparison is independent of the status quo. Theorem 9 - Abstain Safety Under the autonomy criterion, the structural margin forces the system to defer when the observable advantage is insufficient to overcome the utility uncertainty represented by |sigma|. The harm parameter delta directly affects this structural margin, allowing the gate to encode different risk tolerances through explicit utility parameters rather than heuristic confidence thresholds. STRATEGIC EXTENSION - NEW IN VERSION 4 Version 4 studies the autonomy rule as a Stackelberg interaction. The defender commits to a decision rule, while a strategic follower observes that rule and may exert effort through evasion, feature inflation, threshold gaming, or another adaptive response. As a result, the observed margins p_a = p_a(tau) and p_h = p_h(tau) may themselves depend on the deployed threshold. Importantly, sigma does not change strategically; it remains a utility constant. What changes are the observed margins p_a and p_h, together with the effective uncertainty created by strategic adaptation. Two strategic directions are distinguished. Degradation of observable performance tends to close the autonomy gate and therefore pushes the system toward abstention. Inflation is potentially more dangerous: observable performance may improve while underlying true risk remains unchanged, potentially causing a stale rule to grant autonomy incorrectly. Theorem 10 - Abstain Robustness Under Strategic Adaptation Version 4 compares two regimes. In naive Version 3 deployment, the autonomy gate continues using pre-deployment or stale parameters after the environment adapts. In robust Version 4 deployment, the relevant margins are re-estimated after strategic response and the rule is iterated toward a fixed point. Computational experiments show that stale-parameter deployment can produce non-zero active error under strategic adaptation, whereas re-estimation restores zero active error in the tested configurations. Monte Carlo verification produced the following results: sigma = 20: naive 9/20 active errors, robust 0; sigma = 0: naive 18/20 active errors, robust 0; second sigma = 0 configuration: naive 17/20 active errors, robust 0; and second sigma = 20 configuration: naive 12/20 active errors, robust 0. For the wide-margin sigma = 60 and sigma = 180 configurations, the gate remains closed from the outset. In these tested regimes, strategic feedback therefore cannot overturn the autonomy decision. Theorem 11 - Strategic Stability of Abstention If the rule enters the ABSTAIN state, strategic effort directed specifically at the autonomous agent’s threshold loses its payoff. Let adversarial utility be U(e) = P(agent decision passes | e) * V - c * e. Under ABSTAIN, the autonomous agent does not produce the final decision, so P(agent decision passes | e) = 0. Therefore U(e | ABSTAIN) = -c * e, which is maximized at e* = 0. Thus, abstention is not only a safety state: under the stated strategic model, it also removes the incentive to exert effort against the autonomous threshold. This result does not require the monotonicity assumption used in the convergence analysis. Theorem 12 - Convergence of Re-estimation Under the stated monotonic response assumptions, Version 4 studies the iterative rule tau_(t+1) = F(tau_t). The computational implementation uses quantile position q as the state, with F(q) representing the new distributional quantile position associated with a fixed risk boundary. Two structurally different response families were evaluated: linear and logistic/saturating responses. Across three response scales and five starting conditions for each family, all 30/30 computational runs converged to the same observed fixed point, q* approximately 0.999, with zero active error in the tested runs. The result establishes convergence under the stated monotonic setting; it does not establish general uniqueness of the fixed point. Theorem 13 - Zero Active Error at Equilibrium At a fixed point tau*, the rule is evaluated using the same equilibrium margins generated under that rule. The resulting state therefore either ABSTAINS, producing no autonomous active error, or ACTS when the T8 criterion is satisfied using the equilibrium margins. Under the assumptions of the framework, active error(tau*) = 0. EMPIRICAL VALIDATION AND THE NAIVE FALLACY Healthcare Data - TWINS The original framework was validated on the Louizos et al. TWINS dataset with n = 71,345. The analysis exposes the structural limitation of the naive benefit rate a - b = PNS - P(H). In the high-risk preterm stratum, the naive metric falls to 0.19% while the true benefit rate is 2.7%, producing approximately a 14-fold underestimation because true benefit and active harm cancel each other in the aggregate. Financial Market DGP and yfinance Data The framework was also transferred to trading decisions using synthetic segmentations and approximately three years of public financial-market data across 380 non-overlapping decision units. Observable counterfactual market movements demonstrate how additional counterfactual information can tighten structural uncertainty and allow the benefit interval to bound true utility outside clinical settings. LLM Guardrail Validation - FiFAR / OpenL2D The Version 3 autonomy gate was evaluated using a public fraud dataset with AI decisions and 50 synthetic human analysts, with n = 30,622 cases. The T8/T9 gate correctly executed 54/54 evaluated segment decisions between ACT and ABSTAIN using explicit utility bounds rather than arbitrary confidence thresholds. Version 4 revisits FiFAR specifically as a strategic-feedback stress test. The dataset exhibits a strong baseline imbalance: fraud prevalence is approximately 12%, giving a trivial majority-class accuracy of approximately 88%. In the tested setup, the human analysts are substantially below this level. This creates an informative negative result: the tested strategic perturbation is unable to push the AI below the human comparison boundary, making the autonomy decision structurally robust in this tested regime. This leads to an important empirical scope finding: Goodhart-type strategic feedback becomes consequential primarily near narrow decision margins. Large performance separations or sufficiently wide structural margins can make the autonomy gate robust before iterative correction is required. Lending Club - New Strategic Validation in Version 4 Version 4 introduces a second real-data strategic validation using approximately 1.34 million completed Lending Club loans, with an observed default rate of approximately 20%. Credit quality meaningfully discriminates risk in the evaluated data, with default rates ranging approximately from 26.8% to 11.1% across the relevant credit-quality range. The strategic simulation model

Authors

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-25
DOI
https://doi.org/10.5281/zenodo.22962244
Primary Topic
Ethics and Social Impacts of AI
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
preprint

A Decision-Theoretic Extension of Causal Unit Selection: Identifiability, Threshold Rules, and the Fallacy of Naive Benefit Rates.

Joona Matti Ensio Eskelinen
Zenodo (CERN European Organization for Nuclear Research)
Ethics and Social Impacts of AI
preprint

A Decision-Theoretic Extension of Causal Unit Selection: Identifiability, Threshold Rules, and the Fallacy of Naive Benefit Rates.

Joona Matti Ensio Eskelinen
preprint en

Abstract

ABSTRACT AND OVERVIEW This repository contains the official preprint and Python replication code for the research paper “A Decision-Theoretic Extension of Causal Unit Selection: Identifiability, Threshold Rules, and the Fallacy of Naive Benefit Rates” (Eskelinen, Version 4). The work formalizes and advances the principles of individualized causal decision-making by extending the foundational Li-Pearl (2021) benefit function into a decision-theoretic framework. While standard unit-selection frameworks focus on bounding unobserved counterfactual response types (Benefited, Always-taker, Never-taker, and Harm), this paper derives explicit operational decision rules for group selection under counterfactual uncertainty and finite-data constraints. Version 3 extended these results into an LLM Agent Autonomy Guardrail, deriving a mathematical gate for deciding when an AI agent may act autonomously and when it must defer to a human. Version 4 extends the Version 3 autonomy framework to strategic environments. The central question is what happens after the autonomy rule itself is deployed and becomes observable to the environment. Once agents, applicants, adversaries, or other systems can respond to the rule, the observed success margins used by the autonomy gate may become endogenous to the policy itself. This creates a Lucas-critique / Goodhart-type problem: parameters estimated before deployment need not remain invariant after actors begin optimizing against the deployed decision rule. Version 4 does not replace or reinvent the mathematics of T5-T9. Instead, it develops a strategic extension of the existing causal unit-selection and abstention structure and studies when re-estimation restores the safety properties of the original autonomy gate. CORE MATHEMATICAL AND THEORETICAL CONTRIBUTIONS Theorem 5 - Identifiability The target benefit function f(c) is point-identifiable purely from observable causal effects if and only if the structural master coefficient sigma = beta - gamma - theta + delta equals zero. When sigma is non-zero, counterfactual shifts can alter the true benefit function while leaving observable causal margins unchanged. Theorem 6 - Decision Criterion and the Abstain Rule Under fundamental counterfactual uncertainty where the Probability of Necessity and Sufficiency (PNS) is unknown, target groups c0 and c1 are strictly decidable if and only if |W1 - W0| > |sigma|. When this condition fails, the structural intervals overlap and the rule abstains from selection. Theorem 7 - Reliability Under Estimation For a true gap margin D = |W1 - W0| - |sigma| > 0, the decision rule yields zero active errors asymptotically, with the required empirical sample size scaling proportionally to 1 / D^2. When D <= 0, additional data alone cannot resolve the underlying structural uncertainty. Theorem 8 - Autonomy Criterion Extending T5-T7 to agent environments, an agent’s action is strictly and decidably better than deferring to a human baseline if and only if |gamma - delta| * |p_a - p_h| > |sigma|, where p_a denotes the agent success rate and p_h denotes the human success rate. The comparison is independent of the status quo. Theorem 9 - Abstain Safety Under the autonomy criterion, the structural margin forces the system to defer when the observable advantage is insufficient to overcome the utility uncertainty represented by |sigma|. The harm parameter delta directly affects this structural margin, allowing the gate to encode different risk tolerances through explicit utility parameters rather than heuristic confidence thresholds. STRATEGIC EXTENSION - NEW IN VERSION 4 Version 4 studies the autonomy rule as a Stackelberg interaction. The defender commits to a decision rule, while a strategic follower observes that rule and may exert effort through evasion, feature inflation, threshold gaming, or another adaptive response. As a result, the observed margins p_a = p_a(tau) and p_h = p_h(tau) may themselves depend on the deployed threshold. Importantly, sigma does not change strategically; it remains a utility constant. What changes are the observed margins p_a and p_h, together with the effective uncertainty created by strategic adaptation. Two strategic directions are distinguished. Degradation of observable performance tends to close the autonomy gate and therefore pushes the system toward abstention. Inflation is potentially more dangerous: observable performance may improve while underlying true risk remains unchanged, potentially causing a stale rule to grant autonomy incorrectly. Theorem 10 - Abstain Robustness Under Strategic Adaptation Version 4 compares two regimes. In naive Version 3 deployment, the autonomy gate continues using pre-deployment or stale parameters after the environment adapts. In robust Version 4 deployment, the relevant margins are re-estimated after strategic response and the rule is iterated toward a fixed point. Computational experiments show that stale-parameter deployment can produce non-zero active error under strategic adaptation, whereas re-estimation restores zero active error in the tested configurations. Monte Carlo verification produced the following results: sigma = 20: naive 9/20 active errors, robust 0; sigma = 0: naive 18/20 active errors, robust 0; second sigma = 0 configuration: naive 17/20 active errors, robust 0; and second sigma = 20 configuration: naive 12/20 active errors, robust 0. For the wide-margin sigma = 60 and sigma = 180 configurations, the gate remains closed from the outset. In these tested regimes, strategic feedback therefore cannot overturn the autonomy decision. Theorem 11 - Strategic Stability of Abstention If the rule enters the ABSTAIN state, strategic effort directed specifically at the autonomous agent’s threshold loses its payoff. Let adversarial utility be U(e) = P(agent decision passes | e) * V - c * e. Under ABSTAIN, the autonomous agent does not produce the final decision, so P(agent decision passes | e) = 0. Therefore U(e | ABSTAIN) = -c * e, which is maximized at e* = 0. Thus, abstention is not only a safety state: under the stated strategic model, it also removes the incentive to exert effort against the autonomous threshold. This result does not require the monotonicity assumption used in the convergence analysis. Theorem 12 - Convergence of Re-estimation Under the stated monotonic response assumptions, Version 4 studies the iterative rule tau_(t+1) = F(tau_t). The computational implementation uses quantile position q as the state, with F(q) representing the new distributional quantile position associated with a fixed risk boundary. Two structurally different response families were evaluated: linear and logistic/saturating responses. Across three response scales and five starting conditions for each family, all 30/30 computational runs converged to the same observed fixed point, q* approximately 0.999, with zero active error in the tested runs. The result establishes convergence under the stated monotonic setting; it does not establish general uniqueness of the fixed point. Theorem 13 - Zero Active Error at Equilibrium At a fixed point tau*, the rule is evaluated using the same equilibrium margins generated under that rule. The resulting state therefore either ABSTAINS, producing no autonomous active error, or ACTS when the T8 criterion is satisfied using the equilibrium margins. Under the assumptions of the framework, active error(tau*) = 0. EMPIRICAL VALIDATION AND THE NAIVE FALLACY Healthcare Data - TWINS The original framework was validated on the Louizos et al. TWINS dataset with n = 71,345. The analysis exposes the structural limitation of the naive benefit rate a - b = PNS - P(H). In the high-risk preterm stratum, the naive metric falls to 0.19% while the true benefit rate is 2.7%, producing approximately a 14-fold underestimation because true benefit and active harm cancel each other in the aggregate. Financial Market DGP and yfinance Data The framework was also transferred to trading decisions using synthetic segmentations and approximately three years of public financial-market data across 380 non-overlapping decision units. Observable counterfactual market movements demonstrate how additional counterfactual information can tighten structural uncertainty and allow the benefit interval to bound true utility outside clinical settings. LLM Guardrail Validation - FiFAR / OpenL2D The Version 3 autonomy gate was evaluated using a public fraud dataset with AI decisions and 50 synthetic human analysts, with n = 30,622 cases. The T8/T9 gate correctly executed 54/54 evaluated segment decisions between ACT and ABSTAIN using explicit utility bounds rather than arbitrary confidence thresholds. Version 4 revisits FiFAR specifically as a strategic-feedback stress test. The dataset exhibits a strong baseline imbalance: fraud prevalence is approximately 12%, giving a trivial majority-class accuracy of approximately 88%. In the tested setup, the human analysts are substantially below this level. This creates an informative negative result: the tested strategic perturbation is unable to push the AI below the human comparison boundary, making the autonomy decision structurally robust in this tested regime. This leads to an important empirical scope finding: Goodhart-type strategic feedback becomes consequential primarily near narrow decision margins. Large performance separations or sufficiently wide structural margins can make the autonomy gate robust before iterative correction is required. Lending Club - New Strategic Validation in Version 4 Version 4 introduces a second real-data strategic validation using approximately 1.34 million completed Lending Club loans, with an observed default rate of approximately 20%. Credit quality meaningfully discriminates risk in the evaluated data, with default rates ranging approximately from 26.8% to 11.1% across the relevant credit-quality range. The strategic simulation model

Zenodo (CERN European Organization for Nuclear Research)
Peace, Justice and strong institutions
Ethics and Social Impacts of AI
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.