Measuring Consensus Across Language Models and the Limits of Normative Interpretation
Version 2.2.1. A pre-registered relational measurement of where large language models agree. Three registered hypotheses on seven deployments across five vendor families are supported: all-model consensus is rarer than a pooled-marginal independence baseline predicts; it concentrates where one option instantiates a codified norm (85.7% against 45.0%); and the classification replicates on fresh sessions. This version adds a second pre-registered study comparing pre-trained and instruction-tuned weights in four open-weight pairs, four robustness analyses of the norm labels, and an exploratory reanalysis separating answer concentration from directional agreement between models. What changed since version 1.0, and why the change matters: several claims of the first version are withdrawn here. The consensus core is present in published pre-trained weights before instruction tuning, so it is described as machine orthodoxy rather than doxa, and as inherited from a written record that vendor policy also restates rather than as installed by alignment. Pre-trained models find the codified items more determinate than the open items, so normativity and answerability are not identified apart and the normative reading of the main gradient is supported but not established. The two additional main tests do not pass Holm correction at .05. The thirty open-dilemma items on which all seven models agree were authored with no designated option and were never checked against policy text; they are candidates for machine doxa, not a demonstration of it. Data, code and verification records for the reanalysis are archived separately at 10.5281/zenodo.22867323. Files: manuscript, supplement, and title page carrying the declarations. Adversarial review was performed by commissioned model instances within the same AI-assisted workflow and is not external peer review.Note on the title: version 1.0 circulated as “Machine Doxa: Where All Models Agree”. The present version is titled “Measuring Consensus Across Language Models and the Limits of Normative Interpretation”, reflecting the narrower claims this version makes.
Authors
- Toeda Taiko (ORCID: https://orcid.org/0009-0001-7267-0201)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-21
- DOI
- https://doi.org/10.5281/zenodo.22867486
- Primary Topic
- Explainable Artificial Intelligence (XAI)
- Type
- preprint