Coverage-Free Fuzzing: LLM-Guided Refinement of Grammar-Based Test Generators
Property-based testing tools such as Hypothesis can turn a formal grammar into a generator of syntactically valid inputs, but the generator itself is typically static: hand-tuned once and reused unchanged for the rest of a fuzzing campaign. We study whether a large language model (LLM), given only a formal grammar and parser-level feedback – with no coverage instrumentation – can iteratively refine such a generator. We build a fully black-box pipeline targeting the cJSON parser (pinned at v1.7.19): a baseline Hypothesis strategy derived directly from an ANTLR JSON grammar generates inputs, a sanitizer-instrumented harness classifies each as accepted, rejected, or crashing, and a refinement loop feeds an LLM a summary of three observable proxy signals – acceptance rate, the structural diversity of the shapes produced, and the distinct rejection signatures seen – asking it to propose a revised generator. Every proposal is statically sandboxed (AST-checked, import-restricted) and validated before it ever touches the target binary. Across fifteen independent, seeded runs of up to five refinement iterations each, LLM-guided refinement raises mean acceptance rate from 58.2% to 97.1% relative to a static baseline, using no coverage data at all. Structural fingerprints and rejection signatures are reported as descriptive signals; because the refined arm executes more iterations, their aggregate counts are not treated as like-for-like effects. A follow-up ablation over which subset of this feedback the proposer sees – counts alone, counts with rejection signatures, or the full signal – finds acceptance-rate gains are broadly similar across the three conditions, with full feedback preserving more rejection-signature diversity than counts alone (four usable runs per condition; reported descriptively, not tested). A pre-registered replication on a second parser, parson, with the same proposer and design finds the opposite: refinement lowers mean acceptance from 95.6% to 92.4% (p ≈ 0.024), and no crash occurs in 25,500 sanitizer-instrumented executions. Every rejected input is explained: parson accepts any bytes after a complete value, so the static baseline starts near its ceiling, while the refined generators produce grammar-valid constructs that parson refuses without reporting why. The acceptance proxy is thus only as useful as the target is strict and informative about rejections. We discuss how grammar-implementation mismatches – for instance, duplicate JSON object keys accepted by one parser and rejected by the other under an identical grammar – are themselves a source of fuzzing-relevant signal that a purely formal grammar cannot supply, and we release the full pipeline, harnesses, and reproducibility tooling.
Authors
- Ziyad Mohammad Mansy Ibrahim (ORCID: https://orcid.org/0009-0008-3499-3828)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-29
- DOI
- https://doi.org/10.5281/zenodo.23026816
- Primary Topic
- Software Testing and Debugging Techniques
- Type
- preprint