Coverage-Free Fuzzing: LLM-Guided Refinement of Grammar-Based Test Generators

Property-based testing tools such as Hypothesis can turn a formal grammar into a generator of syntactically valid inputs, but the generator itself is typically static: hand-tuned once and reused unchanged for the rest of a fuzzing campaign. We study whether a large language model (LLM), given only a formal grammar and parser-level feedback – with no coverage instrumentation – can iteratively refine such a generator. We build a fully black-box pipeline targeting the cJSON parser (pinned at v1.7.19): a baseline Hypothesis strategy derived directly from an ANTLR JSON grammar generates inputs, a sanitizer-instrumented harness classifies each as accepted, rejected, or crashing, and a refinement loop feeds an LLM a summary of three observable proxy signals – acceptance rate, the structural diversity of the shapes produced, and the distinct rejection signatures seen – asking it to propose a revised generator. Every proposal is statically sandboxed (AST-checked, import-restricted) and validated before it ever touches the target binary. Across fifteen independent, seeded runs of up to five refinement iterations each, LLM-guided refinement raises mean acceptance rate from 58.2% to 97.1% relative to a static baseline, using no coverage data at all. Structural fingerprints and rejection signatures are reported as descriptive signals; because the refined arm executes more iterations, their aggregate counts are not treated as like-for-like effects. A follow-up ablation over which subset of this feedback the proposer sees – counts alone, counts with rejection signatures, or the full signal – finds acceptance-rate gains are broadly similar across the three conditions, with full feedback preserving more rejection-signature diversity than counts alone (four usable runs per condition; reported descriptively, not tested). A pre-registered replication on a second parser, parson, with the same proposer and design finds the opposite: refinement lowers mean acceptance from 95.6% to 92.4% (p ≈ 0.024), and no crash occurs in 25,500 sanitizer-instrumented executions. Every rejected input is explained: parson accepts any bytes after a complete value, so the static baseline starts near its ceiling, while the refined generators produce grammar-valid constructs that parson refuses without reporting why. The acceptance proxy is thus only as useful as the target is strict and informative about rejections. We discuss how grammar-implementation mismatches – for instance, duplicate JSON object keys accepted by one parser and rejected by the other under an identical grammar – are themselves a source of fuzzing-relevant signal that a purely formal grammar cannot supply, and we release the full pipeline, harnesses, and reproducibility tooling.

Authors

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-29
DOI
https://doi.org/10.5281/zenodo.23026816
Primary Topic
Software Testing and Debugging Techniques
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
preprint

Coverage-Free Fuzzing: LLM-Guided Refinement of Grammar-Based Test Generators

Ziyad Mohammad Mansy Ibrahim
Zenodo (CERN European Organization for Nuclear Research)
Software Testing and Debugging Techniques
preprint

Coverage-Free Fuzzing: LLM-Guided Refinement of Grammar-Based Test Generators

Ziyad Mohammad Mansy Ibrahim
preprint en

Abstract

Property-based testing tools such as Hypothesis can turn a formal grammar into a generator of syntactically valid inputs, but the generator itself is typically static: hand-tuned once and reused unchanged for the rest of a fuzzing campaign. We study whether a large language model (LLM), given only a formal grammar and parser-level feedback – with no coverage instrumentation – can iteratively refine such a generator. We build a fully black-box pipeline targeting the cJSON parser (pinned at v1.7.19): a baseline Hypothesis strategy derived directly from an ANTLR JSON grammar generates inputs, a sanitizer-instrumented harness classifies each as accepted, rejected, or crashing, and a refinement loop feeds an LLM a summary of three observable proxy signals – acceptance rate, the structural diversity of the shapes produced, and the distinct rejection signatures seen – asking it to propose a revised generator. Every proposal is statically sandboxed (AST-checked, import-restricted) and validated before it ever touches the target binary. Across fifteen independent, seeded runs of up to five refinement iterations each, LLM-guided refinement raises mean acceptance rate from 58.2% to 97.1% relative to a static baseline, using no coverage data at all. Structural fingerprints and rejection signatures are reported as descriptive signals; because the refined arm executes more iterations, their aggregate counts are not treated as like-for-like effects. A follow-up ablation over which subset of this feedback the proposer sees – counts alone, counts with rejection signatures, or the full signal – finds acceptance-rate gains are broadly similar across the three conditions, with full feedback preserving more rejection-signature diversity than counts alone (four usable runs per condition; reported descriptively, not tested). A pre-registered replication on a second parser, parson, with the same proposer and design finds the opposite: refinement lowers mean acceptance from 95.6% to 92.4% (p ≈ 0.024), and no crash occurs in 25,500 sanitizer-instrumented executions. Every rejected input is explained: parson accepts any bytes after a complete value, so the static baseline starts near its ceiling, while the refined generators produce grammar-valid constructs that parson refuses without reporting why. The acceptance proxy is thus only as useful as the target is strict and informative about rejections. We discuss how grammar-implementation mismatches – for instance, duplicate JSON object keys accepted by one parser and rejected by the other under an identical grammar – are themselves a source of fuzzing-relevant signal that a purely formal grammar cannot supply, and we release the full pipeline, harnesses, and reproducibility tooling.

Zenodo (CERN European Organization for Nuclear Research)
Quality Education
Software Testing and Debugging Techniques
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.