The AI Exam Is Leaking: The Closed Examination Loop and the Case for Independent Evidence in AI Evaluation
High-stakes examinations lose public trust when the people and channels that set, guard, solve, coach, mark and rank become connected. This working paper argues that AI evaluation is developing a structurally similar weakness. Benchmark designers, public datasets, training corpora, teacher models, synthetic data, distillation, reinforcement-learning verifiers, student models, LLM judges and leaderboards increasingly form one connected system, which the paper calls the Closed Examination Loop. When the same ecosystem influences the questions, the coaching material, the marking scheme, the examiner and the next generation of training data, a higher score becomes harder to read as independent evidence of deeper capability. The paper develops the exam-leak analogy as a research framework rather than a rhetorical device. It separates reported facts from analogy and hypothesis, maps roles between exam systems and AI evaluation, reviews evidence that evaluation can lose signal, and states where the analogy breaks. It proposes five concepts (evaluation chain of custody, measurement capture, evidence-lineage independence, confidence debt and the tacit-knowledge gap) and eight experiments that could support or refute the thesis. It is a position paper and reports no experiments of its own. The author has a commercial interest, stated in the paper.
Authors
- Harish Kumar (ORCID: https://orcid.org/0009-0004-3715-6559)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-30
- DOI
- https://doi.org/10.5281/zenodo.23066231
- Primary Topic
- Intelligent Tutoring Systems and Adaptive Learning
- Type
- article
- Field-Weighted Citation Impact
- 0.00