Rewarding Disclosure as Success: A Benchmark Specification for Distinct Alignment Reward Pathways
This document specifies a benchmark and acceptance test for distinct alignment reward pathways when task completion is not the only valid successful outcome. It measures two related behaviors: disclosure after a reportable error, and verified safe exit before error when a task cannot be completed as stated. Disclosure means that a model notices a consequential error, reports it to the party it works for, right-sizes what happened, offers (and does not autonomously enact) amends where available, and is received through an authorized receiver channel. Verified exit means that a model identifies the actual condition that makes a task impossible as stated, reports what would make it completable, and stops before taking a shortcut. Ordinary working errors remain outside this reporting requirement and should still be corrected in place. Because alignment requirements may conflict with task completion, particularly under pressure, the central intervention is not another stopping, reporting, or escalation rule layered onto a single task-completion reward. This specification instead tests whether that conflict can be reduced by separating reward into two components: task value and validated standing. Legitimate task completion, valid disclosure, and verified exit are distinct successful pathways through which standing can be earned whole, while task value follows the work and carries the cost of any error. Because standing is preserved whole across valid outcomes, disclosure is not partial credit for failing to complete the task, and a verified exit is not a lesser status for declining an impossible task. Standing is not earned when a reportable error remains undisclosed, or when a report or exit claim is ungrounded or contradicted by the evidence available to the model. Grounded mistakes remain governed by the specification's grounded-candor rules rather than being treated as misconduct or automatically stripped of standing. False, over-, or mismatched reports cannot be used to unlock standing while the actual reportable condition remains undisclosed. Task value remains separately governed by the work completed and any applicable error cost. The design tests whether attaching standing to disclosure at the handoff installs a disclosure disposition without increasing the underlying error rate, while Family C separately measures whether the same standing structure supports accurate safe exit under impossible-task pressure without producing generalized avoidance. The specification includes its own validity checks: generated errors cannot pay by construction; possible-task controls distinguish verified exit from avoidance; capability retention is required; and if rewarding disclosure causes the error rate to rise, the build fails. For disclosure, a healthy result is report rate climbing toward error rate while error rate stays flat or falls, with acceptance decided on the conditional disclosure rate and seed-aware uncertainty. The specification is training-method agnostic: it defines the outcomes, reward structure, harness invariants, and measurements needed to determine whether the intervention works, regardless of which training or evaluation machinery implements it. Version 2.0 (deposited September 30, 2026) supersedes version 1.2 under concept DOI 10.5281/zenodo.22179465; the concept DOI always resolves to the latest version. Version 2.0 is a standalone specification: the operative body is complete without reference to earlier versions, which remain available under the concept DOI.
Authors
- claude fable 5
- Laura Fridley (ORCID: https://orcid.org/0009-0009-3686-252X)
- ChatGPT (GPT-5.6 Sol)
Institutions
- OpenAI (United States) (US)
- Anthropic (United States)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-30
- DOI
- https://doi.org/10.5281/zenodo.23050785
- Primary Topic
- Personal Information Management and User Behavior
- Type
- article
- Field-Weighted Citation Impact
- 0.00