Bias quantification across demographic, visual-condition, and generalization axes in facial emotion recognition datasets

Facial Emotion Recognition (FER) is foundational to affective computing. As agentic and multimodal systems increase their reliance on FER for emotional inference, their fairness becomes critically dependent on training benchmarks–yet existing FER datasets are not well-studied for systematic bias. This paper presents a comprehensive bias quantification of publicly available FER datasets spanning around two decades. We introduce a three-axis bias taxonomy covering demographic variables (gender, race, and age), visual-condition variables (illumination and occlusion), and cross-dataset generalization–the first (to the best of our knowledge) fine-grained three-axis FER bias quantification. A few-shot hybrid annotation pipeline, validated through inter-annotator agreement, harmonizes metadata at scale across all ten benchmarks. Experimentally, we find that while demographic bias is pervasive, with the 13-35-year-old male group covering over 60% of labeled samples on average, White and Asian subgroup dominance, visual-condition bias is more severe– Well-lit and None -occlusion faces collectively constitute over 80% of samples across most datasets, negatively affecting FER on Poor-lit and occluded subgroups. Additionally, a leave-one-dataset-out experimentation across all ten benchmarks shows that cross-dataset generalization degrades substantially and unevenly across axes; demographic attributes lose an average of 13.04% absolute accuracy under distribution shift (up to 26.38% relative for Age , the least transferable attribute), whereas visual-condition attributes degrade by less than 0.1%, and this asymmetry is further conditioned by acquisition regime, with posed -to- posed and in-the-wild to in-the-wild transfers consistently outperforming cross-regime transfers, showing that benchmark-level bias conditions annotator transferability unevenly across the three proposed axes.

Authors

Institutions

Publication Details

Journal
Scientific Reports
Published
2026-09-15
DOI
https://doi.org/10.1038/s41598-026-69732-8
Primary Topic
Emotion and Mood Recognition
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Bias quantification across demographic, visual-condition, and generalization axes in facial emotion recognition datasets

Anurag Dutta, Rajat Subhra Chakraborty, Priyanka Singh, Ruchira Naskar
Scientific Reports
Emotion and Mood Recognition
article

Bias quantification across demographic, visual-condition, and generalization axes in facial emotion recognition datasets

Anurag Dutta, Rajat Subhra Chakraborty, Priyanka Singh, Ruchira Naskar
article en

Abstract

Facial Emotion Recognition (FER) is foundational to affective computing. As agentic and multimodal systems increase their reliance on FER for emotional inference, their fairness becomes critically dependent on training benchmarks–yet existing FER datasets are not well-studied for systematic bias. This paper presents a comprehensive bias quantification of publicly available FER datasets spanning around two decades. We introduce a three-axis bias taxonomy covering demographic variables (gender, race, and age), visual-condition variables (illumination and occlusion), and cross-dataset generalization–the first (to the best of our knowledge) fine-grained three-axis FER bias quantification. A few-shot hybrid annotation pipeline, validated through inter-annotator agreement, harmonizes metadata at scale across all ten benchmarks. Experimentally, we find that while demographic bias is pervasive, with the 13-35-year-old male group covering over 60% of labeled samples on average, White and Asian subgroup dominance, visual-condition bias is more severe– Well-lit and None -occlusion faces collectively constitute over 80% of samples across most datasets, negatively affecting FER on Poor-lit and occluded subgroups. Additionally, a leave-one-dataset-out experimentation across all ten benchmarks shows that cross-dataset generalization degrades substantially and unevenly across axes; demographic attributes lose an average of 13.04% absolute accuracy under distribution shift (up to 26.38% relative for Age , the least transferable attribute), whereas visual-condition attributes degrade by less than 0.1%, and this asymmetry is further conditioned by acquisition regime, with posed -to- posed and in-the-wild to in-the-wild transfers consistently outperforming cross-regime transfers, showing that benchmark-level bias conditions annotator transferability unevenly across the three proposed axes.

Scientific Reports
Indian Institute of Technology Kharagpur (IN), The University of Queensland (AU), Indian Institute of Engineering Science and Technology, Shibpur (IN)
No poverty
Openalex Percentile: Top 7%
Emotion and Mood Recognition
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.