Bias Mitigation in Federated Learning Through Feature Integration

Machine learning models are trained on datasets where undesired features have spurious correlations with the target variable, resulting in the bias shifts of the model output to a different class. In federated learning (FL), a common problem is data distribution skew. Since each client has a different distribution, models trained on these local datasets are likely to learn different biases. This phenomenon will degrade the performance of the aggregated global model and then prevent the global model from learning actual intrinsic decisions. Consequently, addressing and mitigating these biases in FL is critical to the overall effectiveness of the global model. Most existing works are designed for centralized machine learning using data augmentation or regularization terms. Nevertheless, these methods may not generalize effectively when directly migrated to federated learning scenarios due to the lack of direct access to the external data of other clients. Recent work for debiasing in FL is based on data augmentation in image space, which generates new bias-conflicting images for further training, leading to high costs and limited ability to create sufficiently diverse and realistic samples. To tackle these challenges, we propose a novel feature-level augmentation approach, named FedInt, for debiasing in federated learning. The feature-level augmentation in latent space allows the model to capture and smooth out different directions of the decision boundary, making it possible to synthesize diverse bias-conflicting samples for each client with lower training costs. Experiments on three datasets demonstrate the effectiveness of our proposed method in terms of debiasing performance, and it outperforms traditional data augmentation techniques and various typical FL methods. Our code is available on an anonymous GitHub repository: https://anonymous.4open.science/r/FedInt-C710889456/

Authors

Institutions

Publication Details

Journal
ACM Transactions on Knowledge Discovery from Data
Published
2026-09-30
DOI
https://doi.org/10.1145/3848122
Primary Topic
Privacy-Preserving Technologies in Data
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Bias Mitigation in Federated Learning Through Feature Integration

Kunhao Li, Wanyu Lin, Lei Yang, Hao-Rui Chen et al.
ACM Transactions on Knowledge Discovery from Data
Privacy-Preserving Technologies in Data
article

Bias Mitigation in Federated Learning Through Feature Integration

Kunhao Li, Wanyu Lin, Lei Yang, Hao-Rui Chen, Ziyi Zhang, Mingxuan Ouyang
article en

Abstract

Machine learning models are trained on datasets where undesired features have spurious correlations with the target variable, resulting in the bias shifts of the model output to a different class. In federated learning (FL), a common problem is data distribution skew. Since each client has a different distribution, models trained on these local datasets are likely to learn different biases. This phenomenon will degrade the performance of the aggregated global model and then prevent the global model from learning actual intrinsic decisions. Consequently, addressing and mitigating these biases in FL is critical to the overall effectiveness of the global model. Most existing works are designed for centralized machine learning using data augmentation or regularization terms. Nevertheless, these methods may not generalize effectively when directly migrated to federated learning scenarios due to the lack of direct access to the external data of other clients. Recent work for debiasing in FL is based on data augmentation in image space, which generates new bias-conflicting images for further training, leading to high costs and limited ability to create sufficiently diverse and realistic samples. To tackle these challenges, we propose a novel feature-level augmentation approach, named FedInt, for debiasing in federated learning. The feature-level augmentation in latent space allows the model to capture and smooth out different directions of the decision boundary, making it possible to synthesize diverse bias-conflicting samples for each client with lower training costs. Experiments on three datasets demonstrate the effectiveness of our proposed method in terms of debiasing performance, and it outperforms traditional data augmentation techniques and various typical FL methods. Our code is available on an anonymous GitHub repository: https://anonymous.4open.science/r/FedInt-C710889456/

ACM Transactions on Knowledge Discovery from Data
Hong Kong Polytechnic University (HK), South China University of Technology (CN)
Peace, Justice and strong institutions
Openalex Percentile: Top 9%
Privacy-Preserving Technologies in Data
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.