Average Fairness Gains Do Not Imply Directionally Consistent Intervention Effects in Tabular In-Context Learning: A Reliability Audit of Fair-TabICL

Fair tabular in-context learning can improve group-fairness metrics through context selection without retraining the underlying foundation model. Aggregate improvements, however, do not answer whether a stochastic intervention moves fairness in the intended direction reliably across valid realizations. We audit the paired effect-direction reliability of uncertainty-based context selection in Fair-TabICL as a focused case study on ACSIncome with TabICL. Under the primary shared-priority protocol, uncertainty-based selection has favorable average effects on demographic parity, equality of opportunity, and equalized odds, yet substantial realization-level reversals remain. Two 125-pair repartitioning controls strengthen the average-effect evidence within the evaluated dataset and protocols: hierarchical-bootstrap confidence intervals for all three mean fairness effects remain below zero. The qualitative reliability pattern persists across historical/current TabICL checkpoints, an 8-to-32-estimator control, leakage-free outer-fold scaling, and a matched CPU/GPU backend bridge. When context sampling is decoupled, all three fairness point estimates remain favorable, but variance increases and the corresponding confidence intervals include zero. These results separate two protocol-conditioned properties that are often conflated: favorable average fairness effect and reliable intervention direction. For stochastic context-selection methods, fairness evaluation should report paired effect distributions and direction consistency in addition to aggregate means.

Authors

Institutions

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-15
DOI
https://doi.org/10.5281/zenodo.22765564
Primary Topic
Ethics and Social Impacts of AI
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Average Fairness Gains Do Not Imply Directionally Consistent Intervention Effects in Tabular In-Context Learning: A Reliability Audit of Fair-TabICL

Ran Tao
Zenodo (CERN European Organization for Nuclear Research)
Ethics and Social Impacts of AI
article

Average Fairness Gains Do Not Imply Directionally Consistent Intervention Effects in Tabular In-Context Learning: A Reliability Audit of Fair-TabICL

Ran Tao
article en

Abstract

Fair tabular in-context learning can improve group-fairness metrics through context selection without retraining the underlying foundation model. Aggregate improvements, however, do not answer whether a stochastic intervention moves fairness in the intended direction reliably across valid realizations. We audit the paired effect-direction reliability of uncertainty-based context selection in Fair-TabICL as a focused case study on ACSIncome with TabICL. Under the primary shared-priority protocol, uncertainty-based selection has favorable average effects on demographic parity, equality of opportunity, and equalized odds, yet substantial realization-level reversals remain. Two 125-pair repartitioning controls strengthen the average-effect evidence within the evaluated dataset and protocols: hierarchical-bootstrap confidence intervals for all three mean fairness effects remain below zero. The qualitative reliability pattern persists across historical/current TabICL checkpoints, an 8-to-32-estimator control, leakage-free outer-fold scaling, and a matched CPU/GPU backend bridge. When context sampling is decoupled, all three fairness point estimates remain favorable, but variance increases and the corresponding confidence intervals include zero. These results separate two protocol-conditioned properties that are often conflated: favorable average fairness effect and reliable intervention direction. For stochastic context-selection methods, fairness evaluation should report paired effect distributions and direction consistency in addition to aggregate means.

Zenodo (CERN European Organization for Nuclear Research)
Anhui University (CN)
Peace, Justice and strong institutions, Reduced inequalities
Openalex Percentile: Top 7%
Ethics and Social Impacts of AI
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Average Fairness Gains Do Not Imply Directionally Consistent Intervention Effects in Tabular In-Context Learning: A Reliability Audit of Fair-TabICL — Ran Tao · Zenodo (CERN European Organization for Nuclear Research) (2026) | TGRS Research Map | TGRS