The AI Exam Is Leaking: The Closed Examination Loop and the Case for Independent Evidence in AI Evaluation

High-stakes examinations lose public trust when the people and channels that set, guard, solve, coach, mark and rank become connected. This working paper argues that AI evaluation is developing a structurally similar weakness. Benchmark designers, public datasets, training corpora, teacher models, synthetic data, distillation, reinforcement-learning verifiers, student models, LLM judges and leaderboards increasingly form one connected system, which the paper calls the Closed Examination Loop. When the same ecosystem influences the questions, the coaching material, the marking scheme, the examiner and the next generation of training data, a higher score becomes harder to read as independent evidence of deeper capability. The paper develops the exam-leak analogy as a research framework rather than a rhetorical device. It separates reported facts from analogy and hypothesis, maps roles between exam systems and AI evaluation, reviews evidence that evaluation can lose signal, and states where the analogy breaks. It proposes five concepts (evaluation chain of custody, measurement capture, evidence-lineage independence, confidence debt and the tacit-knowledge gap) and eight experiments that could support or refute the thesis. It is a position paper and reports no experiments of its own. The author has a commercial interest, stated in the paper.

Authors

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-30
DOI
https://doi.org/10.5281/zenodo.23066231
Primary Topic
Intelligent Tutoring Systems and Adaptive Learning
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

The AI Exam Is Leaking: The Closed Examination Loop and the Case for Independent Evidence in AI Evaluation

Harish Kumar
Zenodo (CERN European Organization for Nuclear Research)
Intelligent Tutoring Systems and Adaptive Learning
article

The AI Exam Is Leaking: The Closed Examination Loop and the Case for Independent Evidence in AI Evaluation

Harish Kumar
article en

Abstract

High-stakes examinations lose public trust when the people and channels that set, guard, solve, coach, mark and rank become connected. This working paper argues that AI evaluation is developing a structurally similar weakness. Benchmark designers, public datasets, training corpora, teacher models, synthetic data, distillation, reinforcement-learning verifiers, student models, LLM judges and leaderboards increasingly form one connected system, which the paper calls the Closed Examination Loop. When the same ecosystem influences the questions, the coaching material, the marking scheme, the examiner and the next generation of training data, a higher score becomes harder to read as independent evidence of deeper capability. The paper develops the exam-leak analogy as a research framework rather than a rhetorical device. It separates reported facts from analogy and hypothesis, maps roles between exam systems and AI evaluation, reviews evidence that evaluation can lose signal, and states where the analogy breaks. It proposes five concepts (evaluation chain of custody, measurement capture, evidence-lineage independence, confidence debt and the tacit-knowledge gap) and eight experiments that could support or refute the thesis. It is a position paper and reports no experiments of its own. The author has a commercial interest, stated in the paper.

Zenodo (CERN European Organization for Nuclear Research)
Openalex Percentile: Top 9%
Intelligent Tutoring Systems and Adaptive Learning
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

The AI Exam Is Leaking: The Closed Examination Loop and the Case for Independent Evidence in AI Evaluation — Harish Kumar · Zenodo (CERN European Organization for Nuclear Research) (2026) | TGRS Research Map | TGRS