Annotation-Budget Fairness Reliability: How Much Data Does a Trustworthy Fairness Audit of Hate-Speech Classifiers Require?

Fairness audits of hate-speech classifiers are typically performed at a single annotation budget, leaving a basic question unanswered: how much labeled data does a trustworthy fairness audit require? We introduce ABFR (Annotation-Budget Fairness Reliability), a framework that sweeps the training budget and indexes two reliability thresholds—the Performance Reliability Threshold (PRT), the smallest budget at which model performance stabilizes across runs, and the Fairness Reliability Threshold (FRT), the smallest budget at which a probe-based fairness gap stabilizes. Across seven conditions—five from the HASOC hate-speech shared tasks spanning three languages (English, German, code-mixed Hindi), plus two independent English benchmarks, HateXplain and OLID—and a model ladder from classical TF-IDF logistic regression to a 2025-era multilingual encoder, comprising 34 condition × model cells in total, we find a consistent dissociation. Performance reliability resolves at modest budgets and improves with model capability, whereas fairness reliability does not resolve anywhere in the tested range, up to 90% of the available training pool, in nearly every condition and does not improve with capability. A near-deterministic linear model exhibits comparable fairness instability, showing the effect is not solely an artifact of transformer optimization. A controlled experiment holding the training subsample fixed and varying only the optimization seed shows that both data sampling and optimization stochasticity contribute materially. The single condition in which fairness reliability resolves is not explained by probe count as restricting other conditions to equally few groups does not recover reliability. Instead, it is explained by whether the specific audited groups exhibit stable per-group behavior under subsampling. We further show that probe-based worst-group findings are interpretable only where a model measurably fires on neutral probes, and we report exploratory results on caste, identifying a methodological obstacle: probe-neutrality assumptions do not transfer to identity categories, such as caste, whose mention is itself socially marked.

Authors

Institutions

Publication Details

Journal
Information
Published
2026-09-15
DOI
https://doi.org/10.3390/info17090903
Primary Topic
Hate Speech and Cyberbullying Detection
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Annotation-Budget Fairness Reliability: How Much Data Does a Trustworthy Fairness Audit of Hate-Speech Classifiers Require?

Thomas Mandl, Arjun Mukherjee, Sukomal Pal
Information
Hate Speech and Cyberbullying Detection
article

Annotation-Budget Fairness Reliability: How Much Data Does a Trustworthy Fairness Audit of Hate-Speech Classifiers Require?

Thomas Mandl, Arjun Mukherjee, Sukomal Pal
article en

Abstract

Fairness audits of hate-speech classifiers are typically performed at a single annotation budget, leaving a basic question unanswered: how much labeled data does a trustworthy fairness audit require? We introduce ABFR (Annotation-Budget Fairness Reliability), a framework that sweeps the training budget and indexes two reliability thresholds—the Performance Reliability Threshold (PRT), the smallest budget at which model performance stabilizes across runs, and the Fairness Reliability Threshold (FRT), the smallest budget at which a probe-based fairness gap stabilizes. Across seven conditions—five from the HASOC hate-speech shared tasks spanning three languages (English, German, code-mixed Hindi), plus two independent English benchmarks, HateXplain and OLID—and a model ladder from classical TF-IDF logistic regression to a 2025-era multilingual encoder, comprising 34 condition × model cells in total, we find a consistent dissociation. Performance reliability resolves at modest budgets and improves with model capability, whereas fairness reliability does not resolve anywhere in the tested range, up to 90% of the available training pool, in nearly every condition and does not improve with capability. A near-deterministic linear model exhibits comparable fairness instability, showing the effect is not solely an artifact of transformer optimization. A controlled experiment holding the training subsample fixed and varying only the optimization seed shows that both data sampling and optimization stochasticity contribute materially. The single condition in which fairness reliability resolves is not explained by probe count as restricting other conditions to equally few groups does not recover reliability. Instead, it is explained by whether the specific audited groups exhibit stable per-group behavior under subsampling. We further show that probe-based worst-group findings are interpretable only where a model measurably fires on neutral probes, and we report exploratory results on caste, identifying a methodological obstacle: probe-neutrality assumptions do not transfer to identity categories, such as caste, whose mention is itself socially marked.

InformationVol. 17(9)
University of Hildesheim (DE), Indian Institute of Technology BHU (IN), Banaras Hindu University (IN)
Peace, Justice and strong institutions
Openalex Percentile: Top 8%
Hate Speech and Cyberbullying Detection
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.