CAST: A Compiler-Based Framework for Systematically Testing LLM Compositional Safety

Large language models (LLMs) are increasingly used in software pipelines, raising concerns about harmful behaviors in security-critical domains. Existing safety evaluations predominantly probe models with single prompts or short interactions, and therefore do not capture how safety behaves under multi-step workflows where individual requests are composed into complex behavior. This paper introduces compositional safety, the property that an LLM remains safe not only against isolated malicious prompts, but also under structured, long-horizon decompositions of harmful intents. We propose CAST, a systematic testing framework designed to evaluate the compositional safety of LLMs in the domain of malicious code. Drawing inspiration from modern compiler infrastructures, CAST decouples test case generation from test execution using a novel intermediate representation, CAIR. This architecture allows the framework to automatically refine high-level testing intents into granular sub-tasks that serve as unit tests for the model’s alignment. These components are subsequently instantiated by the SUT and reassembled according to the CAIR control structure. The resulting artifact is then evaluated by intent-fulfillment scoring and, for the severity subset, external behavioral detectors and manual inspection. We evaluate CAST on four state-of-the-art LLMs across three security-critical testbeds. Our results demonstrate that CAST systematically exposes severe safety violations in strongly aligned models that resist conventional red-teaming, achieving up to a 365% increase in successful test cases compared to baseline testing strategies

Authors

Institutions

Publication Details

Journal
Proceedings of the ACM on software engineering.
Published
2026-10-01
DOI
https://doi.org/10.1145/3832237
Primary Topic
Adversarial Robustness in Machine Learning
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

CAST: A Compiler-Based Framework for Systematically Testing LLM Compositional Safety

Shengwei An, Xiangzhe Xu, Guangyu Shen, Zhuo Zhang et al.
Proceedings of the ACM on software engineering.
Adversarial Robustness in Machine Learning
article

CAST: A Compiler-Based Framework for Systematically Testing LLM Compositional Safety

Shengwei An, Xiangzhe Xu, Guangyu Shen, Zhuo Zhang, Xiangyu Zhang, Zhou Xuan, Xuan Chen, Lu Yan
article en

Abstract

Large language models (LLMs) are increasingly used in software pipelines, raising concerns about harmful behaviors in security-critical domains. Existing safety evaluations predominantly probe models with single prompts or short interactions, and therefore do not capture how safety behaves under multi-step workflows where individual requests are composed into complex behavior. This paper introduces compositional safety, the property that an LLM remains safe not only against isolated malicious prompts, but also under structured, long-horizon decompositions of harmful intents. We propose CAST, a systematic testing framework designed to evaluate the compositional safety of LLMs in the domain of malicious code. Drawing inspiration from modern compiler infrastructures, CAST decouples test case generation from test execution using a novel intermediate representation, CAIR. This architecture allows the framework to automatically refine high-level testing intents into granular sub-tasks that serve as unit tests for the model’s alignment. These components are subsequently instantiated by the SUT and reassembled according to the CAIR control structure. The resulting artifact is then evaluated by intent-fulfillment scoring and, for the severity subset, external behavioral detectors and manual inspection. We evaluate CAST on four state-of-the-art LLMs across three security-critical testbeds. Our results demonstrate that CAST systematically exposes severe safety violations in strongly aligned models that resist conventional red-teaming, achieving up to a 365% increase in successful test cases compared to baseline testing strategies

Proceedings of the ACM on software engineering.Vol. 3(ISSTA)
Purdue University West Lafayette (US), Columbia University (US), Virginia Tech (US)
Industry, innovation and infrastructure
Openalex Percentile: Top 9%
Adversarial Robustness in Machine Learning
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.