When the Reviewer Is a Machine: Assessment Without Assessors and the Limits of Automated Judgment

The question usually asked about machine-generated peer review is whether it is good. The answer turns out to be that it is good at some things and not others, and that the distinction between them is not a matter of degree. This article argues that assessment is not a single act but a sequence of seven, and that machine systems perform the first four of them well, the fifth unreliably, the sixth in form only, and the seventh not at all. The seven are comprehension, situating, verification, detection, calibration, verdict, and standing. The first four are recognition acts: they establish what a submission claims, where it sits, whether its claims hold, and what is wrong with it. The fifth and sixth are normative: they weigh what has been found against a standard and commit to a decision. The seventh is not an act performed at a moment but a condition sustained afterwards — remaining available to defend, revise, or be corrected. The decomposition is not merely analytic. It predicts a pattern that two independent studies report and neither names. Machine reviewers identify substantially the same faults as human reviewers, with reported content overlap comparable to the overlap between two humans, while systematically inflating ratings for weaker submissions and aligning with human verdicts most closely on stronger ones. The system names the flaws and does not act on them. This article calls the pattern Detection-Verdict Divergence and argues that it is the signature of a process performing recognition acts without performing normative ones. Two further constructs follow. Simulated Assessment is output bearing the form of an assessment produced without the acts that constitute one; it is not detectable from the artefact, because the artefact is what a real assessment also produces. The Judgment Residue is what remains when every automatable component has been automated: the commitment of a party who can be asked why, who can revise, and who bears the consequence of being wrong. The residue is not a capability gap that better systems will close. It is a relation, and relations are not computed. The article compares four deployment configurations, examines three earlier delegations of judgment in laboratory medicine, credit assessment, and legal disclosure, extends the analysis to regulatory inspection, clinical appraisal, underwriting, and content moderation, and proposes an Assessment Decomposition Audit for venues. It argues that the useful policy question is not whether models may be used but which of the seven components a venue permits them to perform. Paper 3 of 10 in The Answerability Series. Manuscript ID MACH-ASSESS-2026-03. 45 pages, 9 figures, 24 tables, five appendices including an audit worksheet, a component reference card, and a worked contrast between a genuine and a simulated assessment of the same submission — all released under CC BY 4.0.

Authors

Institutions

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-15
DOI
https://doi.org/10.5281/zenodo.22771352
Primary Topic
Academic integrity and plagiarism
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
preprint

When the Reviewer Is a Machine: Assessment Without Assessors and the Limits of Automated Judgment

Syed Shahzad
Zenodo (CERN European Organization for Nuclear Research)
Academic integrity and plagiarism
preprint

When the Reviewer Is a Machine: Assessment Without Assessors and the Limits of Automated Judgment

Syed Shahzad
preprint en

Abstract

The question usually asked about machine-generated peer review is whether it is good. The answer turns out to be that it is good at some things and not others, and that the distinction between them is not a matter of degree. This article argues that assessment is not a single act but a sequence of seven, and that machine systems perform the first four of them well, the fifth unreliably, the sixth in form only, and the seventh not at all. The seven are comprehension, situating, verification, detection, calibration, verdict, and standing. The first four are recognition acts: they establish what a submission claims, where it sits, whether its claims hold, and what is wrong with it. The fifth and sixth are normative: they weigh what has been found against a standard and commit to a decision. The seventh is not an act performed at a moment but a condition sustained afterwards — remaining available to defend, revise, or be corrected. The decomposition is not merely analytic. It predicts a pattern that two independent studies report and neither names. Machine reviewers identify substantially the same faults as human reviewers, with reported content overlap comparable to the overlap between two humans, while systematically inflating ratings for weaker submissions and aligning with human verdicts most closely on stronger ones. The system names the flaws and does not act on them. This article calls the pattern Detection-Verdict Divergence and argues that it is the signature of a process performing recognition acts without performing normative ones. Two further constructs follow. Simulated Assessment is output bearing the form of an assessment produced without the acts that constitute one; it is not detectable from the artefact, because the artefact is what a real assessment also produces. The Judgment Residue is what remains when every automatable component has been automated: the commitment of a party who can be asked why, who can revise, and who bears the consequence of being wrong. The residue is not a capability gap that better systems will close. It is a relation, and relations are not computed. The article compares four deployment configurations, examines three earlier delegations of judgment in laboratory medicine, credit assessment, and legal disclosure, extends the analysis to regulatory inspection, clinical appraisal, underwriting, and content moderation, and proposes an Assessment Decomposition Audit for venues. It argues that the useful policy question is not whether models may be used but which of the seven components a venue permits them to perform. Paper 3 of 10 in The Answerability Series. Manuscript ID MACH-ASSESS-2026-03. 45 pages, 9 figures, 24 tables, five appendices including an audit worksheet, a component reference card, and a worked contrast between a genuine and a simulated assessment of the same submission — all released under CC BY 4.0.

Zenodo (CERN European Organization for Nuclear Research)
Sir Syed University of Engineering and Technology (PK)
Peace, Justice and strong institutions
Academic integrity and plagiarism
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.