Recombination or Discovery? A Retrieval-Grounded Novelty Audit of Machine-Generated Research Papers
Fully machine-generated research papers now arrive by the hundreds, but whether they discover anything or merely recombine existing ideas remains unmeasured. We present a retrieval-grounded, protocol-frozen novelty audit of 166 machine-generated papers (the FARS corpus) against 166 topic-matched ICLR 2025 submissions. We decompose each paper into contribution claims along four facets (purpose, mechanism, evaluation, domain; 549 machine and 494 human contributions) and retrieve prior art per facet under per-paper submission-date cutoffs. A pre-registered two-judge protocol with tiebreak arbitration and deterministic state derivation then classifies each contribution as covered, recombination, or facet-novel. Machine contributions are judged facet-novel more often than human ones (56.1% vs. 35.6%). The gap rests mainly on the purpose facet: requiring an uncovered facet other than purpose shrinks it from 20.5 to 4.1 points (95% CI -0.3 to 8.5), and requiring two uncovered facets leaves 5.8 points (2.3 to 9.4). We argue this comparison must not be taken at face value: a standard-anchored adversarial re-audit flags 15-25% of purpose-novel verdicts as potentially covered, and removing the standard anchor drives the refutation rate to 100% on the same items. Repeating the re-audit with three auditor models yields refutation rates from 0% to 100% on the same items, so an LLM re-audit cannot certify these verdicts without external calibration. We therefore treat automated novelty rates as exploratory upper bounds, since a 106-pair gold prior-art audit finds that the deployed retrieval surfaces the known prior art for only 25-29% of pairs per arm. We release a frozen, blinded human-calibration protocol as an artifact for future validation. A companion integrity audit of 306 Agents4Science 2025 submissions finds hard evidence of data fabrication in 0/47 accepted versus 16/197 rejected submissions (one-sided Fisher p=0.029; p=0.072 when only cases flagged by both coders count). The venue's AI reviewers flagged fabrication in 10 of the 16 rejected cases, so the association reflects both detection and honest disclosure.
Authors
- Tao An
Institutions
- Hawaii Pacific University (US)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-30
- DOI
- https://doi.org/10.5281/zenodo.23061518
- Primary Topic
- Ethics and Social Impacts of AI
- Type
- preprint