Annotation-Budget Fairness Reliability: How Much Data Does a Trustworthy Fairness Audit of Hate-Speech Classifiers Require?
Fairness audits of hate-speech classifiers are typically performed at a single annotation budget, leaving a basic question unanswered: how much labeled data does a trustworthy fairness audit require? We introduce ABFR (Annotation-Budget Fairness Reliability), a framework that sweeps the training budget and indexes two reliability thresholds—the Performance Reliability Threshold (PRT), the smallest budget at which model performance stabilizes across runs, and the Fairness Reliability Threshold (FRT), the smallest budget at which a probe-based fairness gap stabilizes. Across seven conditions—five from the HASOC hate-speech shared tasks spanning three languages (English, German, code-mixed Hindi), plus two independent English benchmarks, HateXplain and OLID—and a model ladder from classical TF-IDF logistic regression to a 2025-era multilingual encoder, comprising 34 condition × model cells in total, we find a consistent dissociation. Performance reliability resolves at modest budgets and improves with model capability, whereas fairness reliability does not resolve anywhere in the tested range, up to 90% of the available training pool, in nearly every condition and does not improve with capability. A near-deterministic linear model exhibits comparable fairness instability, showing the effect is not solely an artifact of transformer optimization. A controlled experiment holding the training subsample fixed and varying only the optimization seed shows that both data sampling and optimization stochasticity contribute materially. The single condition in which fairness reliability resolves is not explained by probe count as restricting other conditions to equally few groups does not recover reliability. Instead, it is explained by whether the specific audited groups exhibit stable per-group behavior under subsampling. We further show that probe-based worst-group findings are interpretable only where a model measurably fires on neutral probes, and we report exploratory results on caste, identifying a methodological obstacle: probe-neutrality assumptions do not transfer to identity categories, such as caste, whose mention is itself socially marked.
Authors
- Thomas Mandl (ORCID: https://orcid.org/0000-0002-8398-9699)
- Arjun Mukherjee (ORCID: https://orcid.org/0000-0002-8896-604X)
- Sukomal Pal (ORCID: https://orcid.org/0000-0001-8743-9830)
Institutions
- University of Hildesheim (DE)
- Indian Institute of Technology BHU (IN)
- Banaras Hindu University (IN)
Publication Details
- Journal
- Information
- Published
- 2026-09-15
- DOI
- https://doi.org/10.3390/info17090903
- Primary Topic
- Hate Speech and Cyberbullying Detection
- Type
- article
- Field-Weighted Citation Impact
- 0.00