A Decision-Theoretic Extension of Causal Unit Selection: Identifiability, Threshold Rules, and the Fallacy of Naive Benefit Rates.
Abstract and Overview This repository contains the official preprint and python replication code for the research paper "A Decision-Theoretic Extension of Causal Unit Selection: Identifiability, Threshold Rules, and the Fallacy of Naive Benefit Rates" (Eskelinen, 2026). The work formalizes and advances the principles of individualized causal decision-making by extending the foundational Li-Pearl (2021) benefit function into a rigorous decision-theoretic framework. While standard unit-selection frameworks focus on bounding unobserved counterfactual response types (Benefited, Always-taker, Never-taker, and Harm), this paper derives explicit operational decision rules for group selection under latent confounding, missing counterfactual data, and sample size constraints. Core Mathematical and Theoretical Contributions: Theorem 5 (Identifiability): We provide a strict mathematical proof demonstrating that the target benefit function f(c) is point-identifiable purely from observable causal effects if and only if the structural master coefficient sigma (defined as beta - gamma - theta + delta) equals exactly zero. When sigma is non-zero, counterfactual shifts alter the true function value while holding observable margins constant. Theorem 6 (Decision Criterion and the Abstain Rule): Under fundamental counterfactual uncertainty where the Probability of Necessity and Sufficiency (PNS) is unknown, we prove that target groups c0 and c1 are strictly decidable if and only if the absolute difference between their risk-adjusted weights, |W1 - W0|, strictly exceeds the absolute value of the master coefficient sigma. If this condition is not met, the mathematical intervals overlap, and the optimal action is to completely abstain from selection. Theorem 7 (Reliability Under Estimation): Assuming a true gap margin D = |W1 - W0| - |sigma| > 0, we prove that the decision rule yields zero active errors, where the required empirical sample size scales proportionally to 1/D^2. Below this threshold, additional data yields diminishing returns and cannot resolve structural uncertainty. Empirical Validation and the Naive Fallacy: Validated on the Louizos et al. TWINS dataset (n=71,345), the framework exposes a severe structural flaw in traditional naive benefit rates (a - b = PNS - P(H)). While the naive metric captures 40% of the true benefit in low-risk populations, it collapses to a massive 14-fold underestimation (0.19% vs a true 2.7%) in high-harm strata (such as the preterm segment) because equal probabilities of true benefit and active harm cancel each other out in the aggregate. File Index and Structural Manifest: A_Decision_Theoretic_Extension_of_Causal_Unit_Selection.pdf - Full preprint manuscript containing the complete formal theorems, proofs, algebraic derivations, and empirical setups. unit_selection_theorems.py - Complete replication script containing the algorithmic implementation of the bounding theorems, the dominance threshold verification loop, and the validation protocols executed on the benchmarking TWINS dataset. License: Creative Commons Attribution-NonCommercial 4.0 International (CC BY-NC 4.0)
Authors
- Joona Matti Ensio Eskelinen
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-24
- DOI
- https://doi.org/10.5281/zenodo.22934160
- Primary Topic
- Advanced Causal Inference Techniques
- Type
- preprint