A Shared Moral-Concern Prompt Cannot Be Assumed to Function Equivalently Across Large Language Models: A Preregistered Cross-Model Test of Persona Framing

Cross-model comparisons of moral judgement assume a shared prompt is a shared instrument. We report a preregistered test of Motivated Violation Construction (MVC), on which a loaded actor identity can move ambiguous behaviour into violation, in Claude Opus 4.7, GPT-5.5, Gemini 3.1 Pro, and DeepSeek-V4-Flash, which classified 20 ambiguous-by-design scenarios under a neutral frame and six loaded personas across 11,200 trials. Two models classified almost every scenario as concerning under the neutral frame, leaving little headroom for the predicted increase. One failed the registered test-retest gate, which by post hoc analysis falsely excludes a stable model 83.9% of the time at that baseline, while a non-registered diagnostic finds a retained-model shift the gate cannot see. Three of the four registered MVC predictions were not supported in the retained three-model panel, the fourth indeterminate under an amendment-fixed rule. The post-gateway battery is exploratory by rule, and its three-test adjudication returned a mixed pattern with no adjudicated winner: two tests favoured valence priming under registered failure semantics, one met its MVC-consistent threshold, concentrated in one model. Persona effects reversed sign across models under one prompt, an inverted question format collapsed expressed concern, and one model expressed concern on mundane controls at rates others never approached. Baseline, headroom, retest behaviour, and concern on those controls differed by model at certified standard, position sensitivity differing in an exploratory recomputation, a configuration we call, post hoc, a model's response regime. A forced binary moral-concern prompt cannot be assumed to function equivalently across models.

Authors

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-18
DOI
https://doi.org/10.5281/zenodo.22837759
Primary Topic
Persona Design and Applications
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
preprint

A Shared Moral-Concern Prompt Cannot Be Assumed to Function Equivalently Across Large Language Models: A Preregistered Cross-Model Test of Persona Framing

Emile Boullineau, José Daniel Muñoz Arciniegas
Zenodo (CERN European Organization for Nuclear Research)
Persona Design and Applications
preprint

A Shared Moral-Concern Prompt Cannot Be Assumed to Function Equivalently Across Large Language Models: A Preregistered Cross-Model Test of Persona Framing

Emile Boullineau, José Daniel Muñoz Arciniegas
preprint en

Abstract

Cross-model comparisons of moral judgement assume a shared prompt is a shared instrument. We report a preregistered test of Motivated Violation Construction (MVC), on which a loaded actor identity can move ambiguous behaviour into violation, in Claude Opus 4.7, GPT-5.5, Gemini 3.1 Pro, and DeepSeek-V4-Flash, which classified 20 ambiguous-by-design scenarios under a neutral frame and six loaded personas across 11,200 trials. Two models classified almost every scenario as concerning under the neutral frame, leaving little headroom for the predicted increase. One failed the registered test-retest gate, which by post hoc analysis falsely excludes a stable model 83.9% of the time at that baseline, while a non-registered diagnostic finds a retained-model shift the gate cannot see. Three of the four registered MVC predictions were not supported in the retained three-model panel, the fourth indeterminate under an amendment-fixed rule. The post-gateway battery is exploratory by rule, and its three-test adjudication returned a mixed pattern with no adjudicated winner: two tests favoured valence priming under registered failure semantics, one met its MVC-consistent threshold, concentrated in one model. Persona effects reversed sign across models under one prompt, an inverted question format collapsed expressed concern, and one model expressed concern on mundane controls at rates others never approached. Baseline, headroom, retest behaviour, and concern on those controls differed by model at certified standard, position sensitivity differing in an exploratory recomputation, a configuration we call, post hoc, a model's response regime. A forced binary moral-concern prompt cannot be assumed to function equivalently across models.

Zenodo (CERN European Organization for Nuclear Research)
Peace, Justice and strong institutions
Persona Design and Applications
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

A Shared Moral-Concern Prompt Cannot Be Assumed to Function Equivalently Across Large Language Models: A Preregistered Cross-Model Test of Persona Framing — Emile Boullineau, José Daniel Muñoz Arciniegas · Zenodo (CERN European Organization for Nuclear Research) (2026) | TGRS Research Map | TGRS