Where Does an LLM Fuzzer's Domain Knowledge Come From? Human, Code-Read, and Self-Acquired Knowledge for Differential Testing of Typed Deserializers

An LLM-guided refinement loop for differential testing of Dart JSON deserializers found divergence in 2.9% of generated documents on its own and in 28.8% when given four observations from a human characterization: the bottleneck was domain knowledge. We ask whether the fuzzer can acquire that knowledge itself. Holding the loop fixed and varying only a 1,500-character knowledge slot, a pre-registered study compares no knowledge, the human's observations, a summary by an agent that reads the deserializers' code, and a summary by an agent that runs 20 black-box probe documents. On Dart (n=5 seeds per arm), human knowledge reaches 30.9%, code-read knowledge 11.6% and none 1.5% (Holm-adjusted p ≈ 0.048 for each pair). Unguided probing reaches 10.5%, not detectably above none (p ≈ 0.69): its probes never remove a top-level field or leave the 64-bit integer range, and it succeeds on one seed in five, by accident. In a pre-registered extension, requiring the agent to cover generic test categories raises probing to 20.2% (p ≈ 0.032 against none; 0.16 after correction across six tests), and a larger model to 12.1%; both still miss the integer-overflow pattern their probes never reach. Across arms, which known patterns the loop finds largely follows which facts the knowledge states. On Kotlin (Gson, Moshi, kotlinx.serialization, Jackson), where divergence is common, unguided probing already helps (67.1% against 50.0% over ten seeds, p ≈ 0.007) and systematic probing reaches 81.6%. Self-characterization works where interesting behavior is dense, and needs systematic coverage where it is narrow.

Authors

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-28
DOI
https://doi.org/10.5281/zenodo.23002933
Primary Topic
Software Testing and Debugging Techniques
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
preprint

Where Does an LLM Fuzzer's Domain Knowledge Come From? Human, Code-Read, and Self-Acquired Knowledge for Differential Testing of Typed Deserializers

Ziyad Mohammad Mansy Ibrahim
Zenodo (CERN European Organization for Nuclear Research)
Software Testing and Debugging Techniques
preprint

Where Does an LLM Fuzzer's Domain Knowledge Come From? Human, Code-Read, and Self-Acquired Knowledge for Differential Testing of Typed Deserializers

Ziyad Mohammad Mansy Ibrahim
preprint en

Abstract

An LLM-guided refinement loop for differential testing of Dart JSON deserializers found divergence in 2.9% of generated documents on its own and in 28.8% when given four observations from a human characterization: the bottleneck was domain knowledge. We ask whether the fuzzer can acquire that knowledge itself. Holding the loop fixed and varying only a 1,500-character knowledge slot, a pre-registered study compares no knowledge, the human's observations, a summary by an agent that reads the deserializers' code, and a summary by an agent that runs 20 black-box probe documents. On Dart (n=5 seeds per arm), human knowledge reaches 30.9%, code-read knowledge 11.6% and none 1.5% (Holm-adjusted p ≈ 0.048 for each pair). Unguided probing reaches 10.5%, not detectably above none (p ≈ 0.69): its probes never remove a top-level field or leave the 64-bit integer range, and it succeeds on one seed in five, by accident. In a pre-registered extension, requiring the agent to cover generic test categories raises probing to 20.2% (p ≈ 0.032 against none; 0.16 after correction across six tests), and a larger model to 12.1%; both still miss the integer-overflow pattern their probes never reach. Across arms, which known patterns the loop finds largely follows which facts the knowledge states. On Kotlin (Gson, Moshi, kotlinx.serialization, Jackson), where divergence is common, unguided probing already helps (67.1% against 50.0% over ten seeds, p ≈ 0.007) and systematic probing reaches 81.6%. Self-characterization works where interesting behavior is dense, and needs systematic coverage where it is narrow.

Zenodo (CERN European Organization for Nuclear Research)
Quality Education
Software Testing and Debugging Techniques
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Where Does an LLM Fuzzer's Domain Knowledge Come From? Human, Code-Read, and Self-Acquired Knowledge for Differential Testing of Typed Deserializers — Ziyad Mohammad Mansy Ibrahim · Zenodo (CERN European Organization for Nuclear Research) (2026) | TGRS Research Map | TGRS