When the Validator Enters the Prompt
External evaluators and deterministic validators are increasingly used to guide second-stage model review, but an evaluator result is not automatically safe downstream context. This study tests one narrow intervention on a fixed 48-case discovery set: a common second-stage self-review instruction is run either alone or with the raw current deterministic validation object inserted into the review context. The design contains 48 paired cases, 96 arm executions, 192 model calls, blind machine grading completed before treatment unblind, and a frozen 24-case AI-assisted, author-adjudicated paired review selected from metadata before judge scores. I used GPT-5.6 Sol to assist analysis and articulation in that review, inspected the underlying paired evidence, and adopted each retained judgment as my own; this channel is not independent or unaided human evaluation. A construct-validity audit materially changed the deterministic interpretation: an originally reported aggregate was rejected because the retained row-level adjudication recorded one literal carry-forward component firing on 95 of 96 rows while classifying zero rows as actual state loss under the intended construct. After removing that invalid component, supported deterministic errors were 3/48 without exposure and 2/48 with exposure (2 exposure-better pairs, 45 ties, 1 exposure-worse pair; exact McNemar p=1.0), providing no detectable deterministic advantage on this set. Blind machine-judge quality was similarly near-flat (3.858 vs. 3.899; difference +0.042; paired-bootstrap 95% CI [-0.115, +0.198]). The 24-case author-adjudicated assisted review moved in the opposite direction: 15 cases favored no exposure, 5 favored exposure, and 4 tied; machine-versus-assisted preference agreement was 9/24. The assisted review identified a plausible contamination pattern in the exposure arm: unsupported validator assertions or provenance were sometimes imported into the answer, unrelated state changed, or blocking behavior expanded. The result is bounded to this implementation and fixed discovery set; its main methodological lesson is that evaluator provenance and scorer validity materially affect what an experiment is permitted to claim.
Authors
- Logan Davis (ORCID: https://orcid.org/0009-0006-8244-5610)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-25
- DOI
- https://doi.org/10.5281/zenodo.22955009
- Primary Topic
- Artificial Intelligence in Healthcare and Education
- Type
- preprint