When the Reviewer Is a Machine: Assessment Without Assessors and the Limits of Automated Judgment
The question usually asked about machine-generated peer review is whether it is good. The answer turns out to be that it is good at some things and not others, and that the distinction between them is not a matter of degree. This article argues that assessment is not a single act but a sequence of seven, and that machine systems perform the first four of them well, the fifth unreliably, the sixth in form only, and the seventh not at all. The seven are comprehension, situating, verification, detection, calibration, verdict, and standing. The first four are recognition acts: they establish what a submission claims, where it sits, whether its claims hold, and what is wrong with it. The fifth and sixth are normative: they weigh what has been found against a standard and commit to a decision. The seventh is not an act performed at a moment but a condition sustained afterwards — remaining available to defend, revise, or be corrected. The decomposition is not merely analytic. It predicts a pattern that two independent studies report and neither names. Machine reviewers identify substantially the same faults as human reviewers, with reported content overlap comparable to the overlap between two humans, while systematically inflating ratings for weaker submissions and aligning with human verdicts most closely on stronger ones. The system names the flaws and does not act on them. This article calls the pattern Detection-Verdict Divergence and argues that it is the signature of a process performing recognition acts without performing normative ones. Two further constructs follow. Simulated Assessment is output bearing the form of an assessment produced without the acts that constitute one; it is not detectable from the artefact, because the artefact is what a real assessment also produces. The Judgment Residue is what remains when every automatable component has been automated: the commitment of a party who can be asked why, who can revise, and who bears the consequence of being wrong. The residue is not a capability gap that better systems will close. It is a relation, and relations are not computed. The article compares four deployment configurations, examines three earlier delegations of judgment in laboratory medicine, credit assessment, and legal disclosure, extends the analysis to regulatory inspection, clinical appraisal, underwriting, and content moderation, and proposes an Assessment Decomposition Audit for venues. It argues that the useful policy question is not whether models may be used but which of the seven components a venue permits them to perform. Paper 3 of 10 in The Answerability Series. Manuscript ID MACH-ASSESS-2026-03. 45 pages, 9 figures, 24 tables, five appendices including an audit worksheet, a component reference card, and a worked contrast between a genuine and a simulated assessment of the same submission — all released under CC BY 4.0.
Authors
- Syed Shahzad (ORCID: https://orcid.org/0009-0001-7323-1577)
Institutions
- Sir Syed University of Engineering and Technology (PK)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-15
- DOI
- https://doi.org/10.5281/zenodo.22771352
- Primary Topic
- Academic integrity and plagiarism
- Type
- preprint