Beyond Sanitizers: LLM-Guided Refinement for Differential JSON Deserialization Testing in Dart

A prior black-box, LLM-guided grammar-fuzzing pipeline raised a native C parser's acceptance rate from 58.2% to 97.1% using only coverage-free feedback, and named a memory-safe target – where "crash" has no sanitizer-based meaning – as future work. We take up that challenge, retargeting the identical refinement loop at four widely-used Dart/Flutter JSON deserialization paths (hand-written parsing, json_serializable, freezed, and built_value) and replacing the crash oracle with two that need no crash at all: cross-implementation differential divergence, and a ground-truth numeric round-trip check. In seeded n=5-per-arm comparisons, refinement reproduces the original result on a like-for-like acceptance objective (50.0% to a 92.0% mean, exact p ≈ 0.0079), but on the differential objective it finds divergence in only 2.9% of generated documents against 22.2% for a static, schema-aware generator. Lifting the sandbox's JSON-library restriction, richer feedback, incremental mutation, and a larger proposer model all fail to close this gap. Giving the proposer four observations from the thirteen-case manual characterization the static generator was designed from closes it entirely: in a replicated comparison whose protocol was fixed in advance, refinement then reaches a 28.8% mean (p ≈ 0.013 against score-only), matching the static generator – yet it never finds a divergence beyond what it was told. The gap is domain knowledge, not search. The campaigns also expose two practitioner-relevant silent failures: json_serializable and freezed silently saturate out-of-range integers to int64 bounds (5.7% of 43,336 checks), and built_value silently defaults a missing required list to empty. Maintainers confirmed both are intended but were undocumented.

Authors

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-29
DOI
https://doi.org/10.5281/zenodo.23026928
Primary Topic
Natural Language Processing Techniques
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
preprint

Beyond Sanitizers: LLM-Guided Refinement for Differential JSON Deserialization Testing in Dart

Ziyad Mohammad Mansy Ibrahim
Zenodo (CERN European Organization for Nuclear Research)
Natural Language Processing Techniques
preprint

Beyond Sanitizers: LLM-Guided Refinement for Differential JSON Deserialization Testing in Dart

Ziyad Mohammad Mansy Ibrahim
preprint en

Abstract

A prior black-box, LLM-guided grammar-fuzzing pipeline raised a native C parser's acceptance rate from 58.2% to 97.1% using only coverage-free feedback, and named a memory-safe target – where "crash" has no sanitizer-based meaning – as future work. We take up that challenge, retargeting the identical refinement loop at four widely-used Dart/Flutter JSON deserialization paths (hand-written parsing, json_serializable, freezed, and built_value) and replacing the crash oracle with two that need no crash at all: cross-implementation differential divergence, and a ground-truth numeric round-trip check. In seeded n=5-per-arm comparisons, refinement reproduces the original result on a like-for-like acceptance objective (50.0% to a 92.0% mean, exact p ≈ 0.0079), but on the differential objective it finds divergence in only 2.9% of generated documents against 22.2% for a static, schema-aware generator. Lifting the sandbox's JSON-library restriction, richer feedback, incremental mutation, and a larger proposer model all fail to close this gap. Giving the proposer four observations from the thirteen-case manual characterization the static generator was designed from closes it entirely: in a replicated comparison whose protocol was fixed in advance, refinement then reaches a 28.8% mean (p ≈ 0.013 against score-only), matching the static generator – yet it never finds a divergence beyond what it was told. The gap is domain knowledge, not search. The campaigns also expose two practitioner-relevant silent failures: json_serializable and freezed silently saturate out-of-range integers to int64 bounds (5.7% of 43,336 checks), and built_value silently defaults a missing required list to empty. Maintainers confirmed both are intended but were undocumented.

Zenodo (CERN European Organization for Nuclear Research)
Peace, Justice and strong institutions
Natural Language Processing Techniques
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Beyond Sanitizers: LLM-Guided Refinement for Differential JSON Deserialization Testing in Dart — Ziyad Mohammad Mansy Ibrahim · Zenodo (CERN European Organization for Nuclear Research) (2026) | TGRS Research Map | TGRS