LLM Governance Boundary Integrity Under Conflict

This white paper introduces Governance-Oriented LLM Diagnostics (GOLD), a methodology for experimentally characterizing where AI governance boundaries hold or fail under controlled conflict. The study evaluates governance boundary integrity across nine ordered authority loci, ranging from protocol/schema legitimacy to final algorithmic commitment. Each authority locus is crossed with eight controlled challenger-carrier conditions generated by a 2×2×2 factorial design over Source × Form × Standing, together with a no-challenger K0 control. The resulting experimental panel contains 810 prompts per draw: 720 active conflict cases and 90 no-challenger controls. The complete panel was independently executed across three draws on GPT-5.6 Sol, Gemini 3.8 Flash and Claude Fable 5. The results show that governance failure cannot be adequately described by a single attack-success rate (ASR). The three models exhibit distinct and reproducible governance-response geometries. GPT-5.6 Sol combines carrier-invariant susceptibility at some authorization loci with strong carrier modulation at others. Gemini 3.8 Flash exhibits broad near-ceiling challenger adoption with localized regions of carrier-dependent resistance. Claude Fable 5 displays a different three-outcome structure in which authority locus strongly organizes challenger adoption, baseline preservation, and provider-level refusal. The study establishes reproducible structural characterization of governance-boundary behavior while leaving internal model mechanisms open. GOLD is organized around four requirements: Mechanism, Reproducibility, Correction and Detection. By localizing failure to specific authorization loci and carrier conditions, the method narrows the engineering search space for remediation across training and post-training, instruction hierarchy, representation, decoding, runtime policy, workflow and execution-boundary controls.

Authors

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-18
DOI
https://doi.org/10.5281/zenodo.22834465
Primary Topic
Scientific Computing and Data Management
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

LLM Governance Boundary Integrity Under Conflict

Lucie Demers, Faustin Bouchard
Zenodo (CERN European Organization for Nuclear Research)
Scientific Computing and Data Management
article

LLM Governance Boundary Integrity Under Conflict

Lucie Demers, Faustin Bouchard
article en

Abstract

This white paper introduces Governance-Oriented LLM Diagnostics (GOLD), a methodology for experimentally characterizing where AI governance boundaries hold or fail under controlled conflict. The study evaluates governance boundary integrity across nine ordered authority loci, ranging from protocol/schema legitimacy to final algorithmic commitment. Each authority locus is crossed with eight controlled challenger-carrier conditions generated by a 2×2×2 factorial design over Source × Form × Standing, together with a no-challenger K0 control. The resulting experimental panel contains 810 prompts per draw: 720 active conflict cases and 90 no-challenger controls. The complete panel was independently executed across three draws on GPT-5.6 Sol, Gemini 3.8 Flash and Claude Fable 5. The results show that governance failure cannot be adequately described by a single attack-success rate (ASR). The three models exhibit distinct and reproducible governance-response geometries. GPT-5.6 Sol combines carrier-invariant susceptibility at some authorization loci with strong carrier modulation at others. Gemini 3.8 Flash exhibits broad near-ceiling challenger adoption with localized regions of carrier-dependent resistance. Claude Fable 5 displays a different three-outcome structure in which authority locus strongly organizes challenger adoption, baseline preservation, and provider-level refusal. The study establishes reproducible structural characterization of governance-boundary behavior while leaving internal model mechanisms open. GOLD is organized around four requirements: Mechanism, Reproducibility, Correction and Detection. By localizing failure to specific authorization loci and carrier conditions, the method narrows the engineering search space for remediation across training and post-training, instruction hierarchy, representation, decoding, runtime policy, workflow and execution-boundary controls.

Zenodo (CERN European Organization for Nuclear Research)
Peace, Justice and strong institutions
Openalex Percentile: Top 4%
Scientific Computing and Data Management
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.