Auditing Fairness Reliability in Tabular In-Context Learning: Composition-Matched Controls and Practical Severity

Fairness interventions for tabular in-context learning are often summarized by averagesover sampled demonstration contexts. An average, however, neither establishes whethera fairness conclusion is reliable across admissible context realizations nor shows whetherthe observed shift is specific to the selected examples or reproducible from simpler contextcomposition. We present a paired reliability audit for uncertainty-based context selectionin tabular in-context learning. Across two datasets, two tabular foundation models, and50 repeated realizations per dataset–model setting, we retain realization-level effects, usehierarchical bootstrap uncertainty, evaluate the same contexts across models, and characterizepractical severity with thresholds fixed before the M12 severity analysis. For every uncertaintyselected context U, we additionally construct a random control C that exactly matches thefour joint target-label and sensitive-attribute counts, yielding the decomposition U-V, C-V,and U-C. Adult Income and Diabetes Race exhibit opposite aggregate directions for equalopportunity-family metrics, yet the decomposition is structurally similar: the compositionmatched contrast tracks much of the original shift while the post-matching residual remainssmall and uncertain in both TabICL and TabPFN. Demographic-parity residuals are moredataset dependent. Practical-threshold analysis further shows that effects of material sizeopposite to the aggregate direction remain common. These results motivate treating fairnessclaims about stochastic context interventions as reliability claims that should be auditedwith repeated paired contexts, explicit matched controls, a matched cross-model audit, andpractical-severity summaries.

Authors

Institutions

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-16
DOI
https://doi.org/10.5281/zenodo.22782885
Primary Topic
Ethics and Social Impacts of AI
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
preprint

Auditing Fairness Reliability in Tabular In-Context Learning: Composition-Matched Controls and Practical Severity

Tao Ran
Zenodo (CERN European Organization for Nuclear Research)
Ethics and Social Impacts of AI
preprint

Auditing Fairness Reliability in Tabular In-Context Learning: Composition-Matched Controls and Practical Severity

Tao Ran
preprint en

Abstract

Fairness interventions for tabular in-context learning are often summarized by averagesover sampled demonstration contexts. An average, however, neither establishes whethera fairness conclusion is reliable across admissible context realizations nor shows whetherthe observed shift is specific to the selected examples or reproducible from simpler contextcomposition. We present a paired reliability audit for uncertainty-based context selectionin tabular in-context learning. Across two datasets, two tabular foundation models, and50 repeated realizations per dataset–model setting, we retain realization-level effects, usehierarchical bootstrap uncertainty, evaluate the same contexts across models, and characterizepractical severity with thresholds fixed before the M12 severity analysis. For every uncertaintyselected context U, we additionally construct a random control C that exactly matches thefour joint target-label and sensitive-attribute counts, yielding the decomposition U-V, C-V,and U-C. Adult Income and Diabetes Race exhibit opposite aggregate directions for equalopportunity-family metrics, yet the decomposition is structurally similar: the compositionmatched contrast tracks much of the original shift while the post-matching residual remainssmall and uncertain in both TabICL and TabPFN. Demographic-parity residuals are moredataset dependent. Practical-threshold analysis further shows that effects of material sizeopposite to the aggregate direction remain common. These results motivate treating fairnessclaims about stochastic context interventions as reliability claims that should be auditedwith repeated paired contexts, explicit matched controls, a matched cross-model audit, andpractical-severity summaries.

Zenodo (CERN European Organization for Nuclear Research)
Anhui University (CN)
No poverty
Ethics and Social Impacts of AI
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.