Counting What Inspection Misses: Auditing an LLM-Generated Quiz Bank for an Educational Game
LLMs make it cheap for a small team to write the assessment content of an educational game, and the result usually reads well. We audited the 500-item bilingual multiple-choice bank of a shipped educational RPG, generated by an LLM coding assistant, against a plan fixed before any audit model was run. Factually, the bank is nearly flawless: two independent solver models from other vendors answered 98.8-100% correctly, and adjudicating every disagreement found two ambiguous stems and no wrong keys (0.4%). As a test, it is not: with the question hidden, both models picked the key from the options alone 56-59% of the time in both languages, against 25% chance. Counting explains much of this: keys cluster in one option slot (overall chi-square = 868), are disproportionately the longest option, and sit among distractors of which roughly a fifth to a third can be eliminated without the relevant knowledge, by an LLM judge whose labels we audited. Exploratory analysis found that about 80 items have their answer given away by another item's stem, a defect that exists only across the set. None of these problems is visible one item at a time. We release the data, prompts, model outputs and adjudications, and propose a low-cost audit recipe in which the cheapest steps are the most revealing.
Authors
- Penggan Zhao
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-10-05
- DOI
- https://doi.org/10.5281/zenodo.23150164
- Primary Topic
- Artificial Intelligence in Education
- Type
- preprint