LayerToFair: An Efficient Post-processing Framework for Layer-aware Fairness Repair of Deep Neural Networks
DNNs are increasingly deployed in high-stakes information systems, where fairness is a critical requirement. Existing post-processing methods can repair unfairness without retraining or accessing original data, but they often treat all layers indiscriminately, resulting in high computational cost and suboptimal effectiveness. We investigate the propagation of sensitive information across layers through the lens of Information Bottleneck (IB) theory and provide empirical evidence that sensitive information diminishes progressively with depth, revealing a Sensitive Information Bottleneck Layer in which sensitive information becomes substantially attenuated and remains suppressed downstream. Guided by these insights, we propose LayerToFair, an efficient post-processing framework for layer-aware fairness repair of DNNs. LayerToFair comprises three steps: probe-based Bottleneck Layer localization, key neuron identification, and GRPO-based neuron output scaling. We evaluated LayerToFair across multiple benchmark datasets and fully connected feedforward networks of varying depths and widths, and compared it against three representative baselines. Experimental results show that LayerToFair achieves up to 71% fairness improvement while maintaining at least 98% of the original model performance, and it outperforms the baselines by up to 29%, 63% and 62%, respectively, with the same repair time budget and comparable model performance guarantee. We have released the implementation to facilitate reproducibility and future research.
Authors
- Zibin Zheng (ORCID: https://orcid.org/0000-0001-7872-7718)
- Hui Dou (ORCID: https://orcid.org/0000-0002-9242-1181)
- Yiwen Zhang (ORCID: https://orcid.org/0000-0001-8709-1088)
- Songyang Fan (ORCID: https://orcid.org/0009-0004-6133-823X)
- jiang He (ORCID: https://orcid.org/0009-0008-3228-3937)
- Rongping Shang (ORCID: https://orcid.org/0009-0008-0027-3742)
- Zhibo Yang (ORCID: https://orcid.org/0009-0001-7643-3799)
Institutions
- Anhui University (CN)
- Sun Yat-sen University (CN)
Publication Details
- Journal
- ACM Transactions on Software Engineering and Methodology
- Published
- 2026-10-03
- DOI
- https://doi.org/10.1145/3850154
- Primary Topic
- Adversarial Robustness in Machine Learning
- Type
- article
- Field-Weighted Citation Impact
- 0.00