CID: Clean-Seed-Free Backdoor Defense via Counterfactual Invariance for Neural Code Models

Neural code models automate core software engineering tasks such as code classification and generation, but remain vulnerable to backdoor attacks. Existing defenses struggle with both injection-based and Semantically-Equivalent Transformation (SET)-based triggers and often require a trusted, pre-verified in-distribution clean seed set, which can be costly to obtain in third-party fine-tuning. This paper introduces Counterfactual Invariance-based Defense (CID), a clean-seed-free defense requiring no a priori trusted in-distribution clean data and grounded in operational counterfactual invariance tests. CID exploits asymmetric counterfactual behavior: clean predictions tend to degrade under semantic context corruption, whereas backdoor predictions show larger representation drift when shortcut-carrying structures are neutralized. Accordingly, CID applies two orthogonal probes—a contextual intervention for semantic-sufficiency testing and a gradient-guided structural intervention for representation-drift testing—to extract a high-purity clean seed set directly from a mixed dataset. CID then uses these seeds to calibrate representation-space filtering over the full dataset. Across four software engineering tasks, six trigger instantiations, poisoning rates from 1% to 10%, multiple model architectures, and multilingual code summarization, CID achieves high poison-detection performance relative to evaluated baselines while keeping false positives low in most settings. Clean-only and selected retraining experiments further show conservative benign-data retention, preserved clean-task utility, and reduced residual attack success rate in challenging Defect Detection settings.

Authors

Institutions

Publication Details

Journal
Proceedings of the ACM on software engineering.
Published
2026-10-01
DOI
https://doi.org/10.1145/3832251
Primary Topic
Adversarial Robustness in Machine Learning
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

CID: Clean-Seed-Free Backdoor Defense via Counterfactual Invariance for Neural Code Models

Hai Jin, Junyao Ye, Deqing Zou, Zhen Li et al.
Proceedings of the ACM on software engineering.
Adversarial Robustness in Machine Learning
article

CID: Clean-Seed-Free Backdoor Defense via Counterfactual Invariance for Neural Code Models

Hai Jin, Junyao Ye, Deqing Zou, Zhen Li, Xi Tang, Shulin Li, Shi Liang
article en

Abstract

Neural code models automate core software engineering tasks such as code classification and generation, but remain vulnerable to backdoor attacks. Existing defenses struggle with both injection-based and Semantically-Equivalent Transformation (SET)-based triggers and often require a trusted, pre-verified in-distribution clean seed set, which can be costly to obtain in third-party fine-tuning. This paper introduces Counterfactual Invariance-based Defense (CID), a clean-seed-free defense requiring no a priori trusted in-distribution clean data and grounded in operational counterfactual invariance tests. CID exploits asymmetric counterfactual behavior: clean predictions tend to degrade under semantic context corruption, whereas backdoor predictions show larger representation drift when shortcut-carrying structures are neutralized. Accordingly, CID applies two orthogonal probes—a contextual intervention for semantic-sufficiency testing and a gradient-guided structural intervention for representation-drift testing—to extract a high-purity clean seed set directly from a mixed dataset. CID then uses these seeds to calibrate representation-space filtering over the full dataset. Across four software engineering tasks, six trigger instantiations, poisoning rates from 1% to 10%, multiple model architectures, and multilingual code summarization, CID achieves high poison-detection performance relative to evaluated baselines while keeping false positives low in most settings. Clean-only and selected retraining experiments further show conservative benign-data retention, preserved clean-task utility, and reduced residual attack success rate in challenging Defect Detection settings.

Proceedings of the ACM on software engineering.Vol. 3(ISSTA)
Jinyintan Hospital (CN), Huazhong University of Science and Technology (CN)
Peace, Justice and strong institutions
Openalex Percentile: Top 9%
Adversarial Robustness in Machine Learning
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.