“Wait, My Tool Can’t Do That?” A Fine-Grained Capability Study of Modern Static Taint Analyzers

Static taint analyzers are widely used in software security, yet aggregate vulnerability-detection metrics provide limited insight into the specific language constructs and analysis capabilities that cause tools to succeed or fail. This paper presents xAST , a fine-grained benchmark containing 844 paired capability cases, implemented as 1,825 individual source files across Python, Go, Java, and JavaScript. The benchmark covers 16 capability dimensions and 76 sub-categories, including context, flow, path, field, and element precision, as well as functions, modules, concurrency, expressions, and dynamic features. Each pair contains a positive instance in which an expected source-to-sink flow should be detected and a structurally related negative instance in which the corresponding alarm should be suppressed. We evaluate eight mainstream static taint analyzers under 14 tool-language configurations. The highest pair pass rate is 61.7%, achieved by CodeQL on JavaScript. Across all 2,881 tool-language pair evaluations, 47.5% pass both instances, 36.1% are false-negative-only failures, 14.3% are false-positive-only failures, and 2.1% fail both instances. The aggregate T-instance detection rate is 61.9%, while the aggregate F-instance suppression rate is 83.6%, showing that missed expected flows are the dominant global failure mode, although individual tools exhibit substantially different detection and suppression profiles. Fine-grained analysis further identifies recurring weaknesses in field and element precision, path-feasibility reasoning, transfer-rule coverage, concurrency and asynchronous modeling, advanced language idioms, alias analysis, and module resolution. These results provide a capability-oriented diagnostic complement to application-level and vulnerability-level benchmarks and offer actionable guidance for tool developers, researchers, and practitioners.

Authors

Institutions

Publication Details

Journal
ACM Transactions on Software Engineering and Methodology
Published
2026-10-03
DOI
https://doi.org/10.1145/3849705
Primary Topic
Security and Verification in Computing
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

“Wait, My Tool Can’t Do That?” A Fine-Grained Capability Study of Modern Static Taint Analyzers

Yanjie Zhao, Shenao Wang, Jian Zhao, Kan Yu et al.
ACM Transactions on Software Engineering and Methodology
Security and Verification in Computing
article

“Wait, My Tool Can’t Do That?” A Fine-Grained Capability Study of Modern Static Taint Analyzers

Yanjie Zhao, Shenao Wang, Jian Zhao, Kan Yu, Haoyu Wang, Lizhong Bian, Yayi Wang, Yan Cheng, Junjie He
article en

Abstract

Static taint analyzers are widely used in software security, yet aggregate vulnerability-detection metrics provide limited insight into the specific language constructs and analysis capabilities that cause tools to succeed or fail. This paper presents xAST , a fine-grained benchmark containing 844 paired capability cases, implemented as 1,825 individual source files across Python, Go, Java, and JavaScript. The benchmark covers 16 capability dimensions and 76 sub-categories, including context, flow, path, field, and element precision, as well as functions, modules, concurrency, expressions, and dynamic features. Each pair contains a positive instance in which an expected source-to-sink flow should be detected and a structurally related negative instance in which the corresponding alarm should be suppressed. We evaluate eight mainstream static taint analyzers under 14 tool-language configurations. The highest pair pass rate is 61.7%, achieved by CodeQL on JavaScript. Across all 2,881 tool-language pair evaluations, 47.5% pass both instances, 36.1% are false-negative-only failures, 14.3% are false-positive-only failures, and 2.1% fail both instances. The aggregate T-instance detection rate is 61.9%, while the aggregate F-instance suppression rate is 83.6%, showing that missed expected flows are the dominant global failure mode, although individual tools exhibit substantially different detection and suppression profiles. Fine-grained analysis further identifies recurring weaknesses in field and element precision, path-feasibility reasoning, transfer-rule coverage, concurrency and asynchronous modeling, advanced language idioms, alias analysis, and module resolution. These results provide a capability-oriented diagnostic complement to application-level and vulnerability-level benchmarks and offer actionable guidance for tool developers, researchers, and practitioners.

ACM Transactions on Software Engineering and Methodology
Ant Group (China) (CN), Huazhong University of Science and Technology (CN)
Openalex Percentile: Top 9%
Security and Verification in Computing
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.