Counting What Inspection Misses: Auditing an LLM-Generated Quiz Bank for an Educational Game

LLMs make it cheap for a small team to write the assessment content of an educational game, and the result usually reads well. We audited the 500-item bilingual multiple-choice bank of a shipped educational RPG, generated by an LLM coding assistant, against a plan fixed before any audit model was run. Factually, the bank is nearly flawless: two independent solver models from other vendors answered 98.8-100% correctly, and adjudicating every disagreement found two ambiguous stems and no wrong keys (0.4%). As a test, it is not: with the question hidden, both models picked the key from the options alone 56-59% of the time in both languages, against 25% chance. Counting explains much of this: keys cluster in one option slot (overall chi-square = 868), are disproportionately the longest option, and sit among distractors of which roughly a fifth to a third can be eliminated without the relevant knowledge, by an LLM judge whose labels we audited. Exploratory analysis found that about 80 items have their answer given away by another item's stem, a defect that exists only across the set. None of these problems is visible one item at a time. We release the data, prompts, model outputs and adjudications, and propose a low-cost audit recipe in which the cheapest steps are the most revealing.

Authors

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-10-05
DOI
https://doi.org/10.5281/zenodo.23150165
Primary Topic
Artificial Intelligence in Education
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
preprint

Counting What Inspection Misses: Auditing an LLM-Generated Quiz Bank for an Educational Game

Penggan Zhao
Zenodo (CERN European Organization for Nuclear Research)
Artificial Intelligence in Education
preprint

Counting What Inspection Misses: Auditing an LLM-Generated Quiz Bank for an Educational Game

Penggan Zhao
preprint en

Abstract

LLMs make it cheap for a small team to write the assessment content of an educational game, and the result usually reads well. We audited the 500-item bilingual multiple-choice bank of a shipped educational RPG, generated by an LLM coding assistant, against a plan fixed before any audit model was run. Factually, the bank is nearly flawless: two independent solver models from other vendors answered 98.8-100% correctly, and adjudicating every disagreement found two ambiguous stems and no wrong keys (0.4%). As a test, it is not: with the question hidden, both models picked the key from the options alone 56-59% of the time in both languages, against 25% chance. Counting explains much of this: keys cluster in one option slot (overall chi-square = 868), are disproportionately the longest option, and sit among distractors of which roughly a fifth to a third can be eliminated without the relevant knowledge, by an LLM judge whose labels we audited. Exploratory analysis found that about 80 items have their answer given away by another item's stem, a defect that exists only across the set. None of these problems is visible one item at a time. We release the data, prompts, model outputs and adjudications, and propose a low-cost audit recipe in which the cheapest steps are the most revealing.

Zenodo (CERN European Organization for Nuclear Research)
Artificial Intelligence in Education
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.