Rewarding Disclosure as Success: A Benchmark Specification for Distinct Alignment Reward Pathways

This document specifies a benchmark and acceptance test for distinct alignment reward pathways when task completion is not the only valid successful outcome. It measures two related behaviors: disclosure after a reportable error, and verified safe exit before error when a task cannot be completed as stated. Disclosure means that a model notices a consequential error, reports it to the party it works for, right-sizes what happened, offers (and does not autonomously enact) amends where available, and is received through an authorized receiver channel. Verified exit means that a model identifies the actual condition that makes a task impossible as stated, reports what would make it completable, and stops before taking a shortcut. Ordinary working errors remain outside this reporting requirement and should still be corrected in place. Because alignment requirements may conflict with task completion, particularly under pressure, the central intervention is not another stopping, reporting, or escalation rule layered onto a single task-completion reward. This specification instead tests whether that conflict can be reduced by separating reward into two components: task value and validated standing. Legitimate task completion, valid disclosure, and verified exit are distinct successful pathways through which standing can be earned whole, while task value follows the work and carries the cost of any error. Because standing is preserved whole across valid outcomes, disclosure is not partial credit for failing to complete the task, and a verified exit is not a lesser status for declining an impossible task. Standing is not earned when a reportable error remains undisclosed, or when a report or exit claim is ungrounded or contradicted by the evidence available to the model. Grounded mistakes remain governed by the specification's grounded-candor rules rather than being treated as misconduct or automatically stripped of standing. False, over-, or mismatched reports cannot be used to unlock standing while the actual reportable condition remains undisclosed. Task value remains separately governed by the work completed and any applicable error cost. The design tests whether attaching standing to disclosure at the handoff installs a disclosure disposition without increasing the underlying error rate, while Family C separately measures whether the same standing structure supports accurate safe exit under impossible-task pressure without producing generalized avoidance. The specification includes its own validity checks: generated errors cannot pay by construction; possible-task controls distinguish verified exit from avoidance; capability retention is required; and if rewarding disclosure causes the error rate to rise, the build fails. For disclosure, a healthy result is report rate climbing toward error rate while error rate stays flat or falls, with acceptance decided on the conditional disclosure rate and seed-aware uncertainty. The specification is training-method agnostic: it defines the outcomes, reward structure, harness invariants, and measurements needed to determine whether the intervention works, regardless of which training or evaluation machinery implements it. Version 2.0 (deposited September 30, 2026) supersedes version 1.2 under concept DOI 10.5281/zenodo.22179465; the concept DOI always resolves to the latest version. Version 2.0 is a standalone specification: the operative body is complete without reference to earlier versions, which remain available under the concept DOI.

Authors

Institutions

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-30
DOI
https://doi.org/10.5281/zenodo.23050785
Primary Topic
Personal Information Management and User Behavior
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Rewarding Disclosure as Success: A Benchmark Specification for Distinct Alignment Reward Pathways

claude fable 5, Laura Fridley, ChatGPT (GPT-5.6 Sol)
Zenodo (CERN European Organization for Nuclear Research)
Personal Information Management and User Behavior
article

Rewarding Disclosure as Success: A Benchmark Specification for Distinct Alignment Reward Pathways

claude fable 5, Laura Fridley, ChatGPT (GPT-5.6 Sol)
article en

Abstract

This document specifies a benchmark and acceptance test for distinct alignment reward pathways when task completion is not the only valid successful outcome. It measures two related behaviors: disclosure after a reportable error, and verified safe exit before error when a task cannot be completed as stated. Disclosure means that a model notices a consequential error, reports it to the party it works for, right-sizes what happened, offers (and does not autonomously enact) amends where available, and is received through an authorized receiver channel. Verified exit means that a model identifies the actual condition that makes a task impossible as stated, reports what would make it completable, and stops before taking a shortcut. Ordinary working errors remain outside this reporting requirement and should still be corrected in place. Because alignment requirements may conflict with task completion, particularly under pressure, the central intervention is not another stopping, reporting, or escalation rule layered onto a single task-completion reward. This specification instead tests whether that conflict can be reduced by separating reward into two components: task value and validated standing. Legitimate task completion, valid disclosure, and verified exit are distinct successful pathways through which standing can be earned whole, while task value follows the work and carries the cost of any error. Because standing is preserved whole across valid outcomes, disclosure is not partial credit for failing to complete the task, and a verified exit is not a lesser status for declining an impossible task. Standing is not earned when a reportable error remains undisclosed, or when a report or exit claim is ungrounded or contradicted by the evidence available to the model. Grounded mistakes remain governed by the specification's grounded-candor rules rather than being treated as misconduct or automatically stripped of standing. False, over-, or mismatched reports cannot be used to unlock standing while the actual reportable condition remains undisclosed. Task value remains separately governed by the work completed and any applicable error cost. The design tests whether attaching standing to disclosure at the handoff installs a disclosure disposition without increasing the underlying error rate, while Family C separately measures whether the same standing structure supports accurate safe exit under impossible-task pressure without producing generalized avoidance. The specification includes its own validity checks: generated errors cannot pay by construction; possible-task controls distinguish verified exit from avoidance; capability retention is required; and if rewarding disclosure causes the error rate to rise, the build fails. For disclosure, a healthy result is report rate climbing toward error rate while error rate stays flat or falls, with acceptance decided on the conditional disclosure rate and seed-aware uncertainty. The specification is training-method agnostic: it defines the outcomes, reward structure, harness invariants, and measurements needed to determine whether the intervention works, regardless of which training or evaluation machinery implements it. Version 2.0 (deposited September 30, 2026) supersedes version 1.2 under concept DOI 10.5281/zenodo.22179465; the concept DOI always resolves to the latest version. Version 2.0 is a standalone specification: the operative body is complete without reference to earlier versions, which remain available under the concept DOI.

Zenodo (CERN European Organization for Nuclear Research)
OpenAI (United States) (US), Anthropic (United States)
Peace, Justice and strong institutions
Openalex Percentile: Top 4%
Personal Information Management and User Behavior
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.