Reasoning or Retrieval? A Controlled Evaluation of Citation Metadata Reliability on Post-Cutoff AI Literature
This preprint presents a controlled evaluation of citation metadata reliability for recent AI literature beyond the effective knowledge cutoff of the evaluated language model. The study uses a frozen benchmark of 160 recent AI papers, including a 30-paper pilot set and a 130-paper held-out MAIN evaluation set. Four Gemini 3.5 Flash-Lite experimental conditions are compared: MEMORY_LOW, MEMORY_HIGH, RETRIEVAL_LOW, and RETRIEVAL_HIGH, alongside a deterministic OpenAlex resolver baseline. The primary metric is the CORE Metadata Score (CMS), which combines author F1, publication-year correctness, arXiv-ID correctness, and related bibliographic metadata components. The study also evaluates abstention behavior, unsupported citation metadata generation, and the relative effects of model memory and external retrieval. During development, an implementation error in the initial MAIN run was identified before outcome analysis: benchmark identifiers had been used instead of exact paper titles. All affected outputs were invalidated, the pilot was forensically re-audited, and the MAIN experiment was rerun using the corrected protocol. The final results reported in this preprint are based on the corrected and reproducibility-audited run. Reproducibility artifacts, evaluation code, benchmark materials, and audit documentation are available in the associated software release. Author: Abhishek Anubhab DashAffiliation: Independent ResearcherORCID: 0009-0000-0386-3571
Authors
- Abhishek Anubhab Dash (ORCID: https://orcid.org/0009-0000-0386-3571)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-30
- DOI
- https://doi.org/10.5281/zenodo.23059293
- Primary Topic
- Biomedical Text Mining and Ontologies
- Type
- preprint