CAST: A Compiler-Based Framework for Systematically Testing LLM Compositional Safety
Large language models (LLMs) are increasingly used in software pipelines, raising concerns about harmful behaviors in security-critical domains. Existing safety evaluations predominantly probe models with single prompts or short interactions, and therefore do not capture how safety behaves under multi-step workflows where individual requests are composed into complex behavior. This paper introduces compositional safety, the property that an LLM remains safe not only against isolated malicious prompts, but also under structured, long-horizon decompositions of harmful intents. We propose CAST, a systematic testing framework designed to evaluate the compositional safety of LLMs in the domain of malicious code. Drawing inspiration from modern compiler infrastructures, CAST decouples test case generation from test execution using a novel intermediate representation, CAIR. This architecture allows the framework to automatically refine high-level testing intents into granular sub-tasks that serve as unit tests for the model’s alignment. These components are subsequently instantiated by the SUT and reassembled according to the CAIR control structure. The resulting artifact is then evaluated by intent-fulfillment scoring and, for the severity subset, external behavioral detectors and manual inspection. We evaluate CAST on four state-of-the-art LLMs across three security-critical testbeds. Our results demonstrate that CAST systematically exposes severe safety violations in strongly aligned models that resist conventional red-teaming, achieving up to a 365% increase in successful test cases compared to baseline testing strategies
Authors
- Shengwei An
- Xiangzhe Xu (ORCID: https://orcid.org/0000-0001-6619-781X)
- Guangyu Shen (ORCID: https://orcid.org/0000-0003-1121-3098)
- Zhuo Zhang (ORCID: https://orcid.org/0000-0002-6515-0021)
- Xiangyu Zhang (ORCID: https://orcid.org/0000-0002-9544-2500)
- Zhou Xuan (ORCID: https://orcid.org/0009-0000-5738-885X)
- Xuan Chen (ORCID: https://orcid.org/0009-0003-1161-119X)
- Lu Yan (ORCID: https://orcid.org/0000-0001-6571-8695)
Institutions
- Purdue University West Lafayette (US)
- Columbia University (US)
- Virginia Tech (US)
Publication Details
- Journal
- Proceedings of the ACM on software engineering.
- Published
- 2026-10-01
- DOI
- https://doi.org/10.1145/3832237
- Primary Topic
- Adversarial Robustness in Machine Learning
- Type
- article
- Field-Weighted Citation Impact
- 0.00