Paraphrastic Resistance: Text Fragility Is Not Decision Instability

Structured outputs constrain an AI system to return a valid category, schema, and numeric scores that downstream software can process, but schema adherence is not the same property as decision stability. This paper reports a controlled pilot across three frontier models (Anthropic Sonnet 4.5, Google Gemini 3.1 Flash Lite, and OpenAI GPT-5) on a supplier-payment classification task, using a frozen corpus of one base case, 20 meaning-preserving reformulations, and 20 loaded perturbations (287 cleaned observations). Three findings are supported: schema-valid output can be decisionally unstable, as one model changed its majority decision on 5 of 20 equivalent reformulations; loaded changes did not produce the same operational response across vendors, with five perturbations escalated by none; and repeated identical inputs produced different discrete decisions in one model at a 21.1% pairwise discordance rate. The operational conclusion is narrow but consequential: schema validity does not imply decision stability, and these properties should be evaluated separately. Independent work at ICML 2026 provides complementary evidence for the same monitoring problem from a different experimental direction. The study is a single-domain pilot; cross-domain replication is required before paradigm-level claims.

Authors

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-10-03
DOI
https://doi.org/10.5281/zenodo.23123501
Primary Topic
Ethics and Social Impacts of AI
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
preprint

Paraphrastic Resistance: Text Fragility Is Not Decision Instability

José López López
Zenodo (CERN European Organization for Nuclear Research)
Ethics and Social Impacts of AI
preprint

Paraphrastic Resistance: Text Fragility Is Not Decision Instability

José López López
preprint en

Abstract

Structured outputs constrain an AI system to return a valid category, schema, and numeric scores that downstream software can process, but schema adherence is not the same property as decision stability. This paper reports a controlled pilot across three frontier models (Anthropic Sonnet 4.5, Google Gemini 3.1 Flash Lite, and OpenAI GPT-5) on a supplier-payment classification task, using a frozen corpus of one base case, 20 meaning-preserving reformulations, and 20 loaded perturbations (287 cleaned observations). Three findings are supported: schema-valid output can be decisionally unstable, as one model changed its majority decision on 5 of 20 equivalent reformulations; loaded changes did not produce the same operational response across vendors, with five perturbations escalated by none; and repeated identical inputs produced different discrete decisions in one model at a 21.1% pairwise discordance rate. The operational conclusion is narrow but consequential: schema validity does not imply decision stability, and these properties should be evaluated separately. Independent work at ICML 2026 provides complementary evidence for the same monitoring problem from a different experimental direction. The study is a single-domain pilot; cross-domain replication is required before paradigm-level claims.

Zenodo (CERN European Organization for Nuclear Research)
Ethics and Social Impacts of AI
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Paraphrastic Resistance: Text Fragility Is Not Decision Instability — José López López · Zenodo (CERN European Organization for Nuclear Research) (2026) | TGRS Research Map | TGRS