CID: Clean-Seed-Free Backdoor Defense via Counterfactual Invariance for Neural Code Models
Neural code models automate core software engineering tasks such as code classification and generation, but remain vulnerable to backdoor attacks. Existing defenses struggle with both injection-based and Semantically-Equivalent Transformation (SET)-based triggers and often require a trusted, pre-verified in-distribution clean seed set, which can be costly to obtain in third-party fine-tuning. This paper introduces Counterfactual Invariance-based Defense (CID), a clean-seed-free defense requiring no a priori trusted in-distribution clean data and grounded in operational counterfactual invariance tests. CID exploits asymmetric counterfactual behavior: clean predictions tend to degrade under semantic context corruption, whereas backdoor predictions show larger representation drift when shortcut-carrying structures are neutralized. Accordingly, CID applies two orthogonal probes—a contextual intervention for semantic-sufficiency testing and a gradient-guided structural intervention for representation-drift testing—to extract a high-purity clean seed set directly from a mixed dataset. CID then uses these seeds to calibrate representation-space filtering over the full dataset. Across four software engineering tasks, six trigger instantiations, poisoning rates from 1% to 10%, multiple model architectures, and multilingual code summarization, CID achieves high poison-detection performance relative to evaluated baselines while keeping false positives low in most settings. Clean-only and selected retraining experiments further show conservative benign-data retention, preserved clean-task utility, and reduced residual attack success rate in challenging Defect Detection settings.
Authors
- Hai Jin (ORCID: https://orcid.org/0000-0002-3934-7605)
- Junyao Ye (ORCID: https://orcid.org/0009-0007-8337-038X)
- Deqing Zou (ORCID: https://orcid.org/0000-0001-8534-5048)
- Zhen Li (ORCID: https://orcid.org/0000-0002-0001-2998)
- Xi Tang (ORCID: https://orcid.org/0009-0006-1925-4686)
- Shulin Li (ORCID: https://orcid.org/0009-0008-8981-0099)
- Shi Liang (ORCID: https://orcid.org/0000-0002-5296-1941)
Institutions
- Jinyintan Hospital (CN)
- Huazhong University of Science and Technology (CN)
Publication Details
- Journal
- Proceedings of the ACM on software engineering.
- Published
- 2026-10-01
- DOI
- https://doi.org/10.1145/3832251
- Primary Topic
- Adversarial Robustness in Machine Learning
- Type
- article
- Field-Weighted Citation Impact
- 0.00