Open-Weight Safety After Loss of Weight Custody: A Framework for Verification, Regulatory Credit, and Residual Risk
Once weights are released, the producer no longer controls the checkpoint. A recipient can copy the weights, modify them, and run the result outside any environment the producer can see. This note argues that the useful question is not whether open weights can be made controllable after that loss of custody. They cannot, in the general case. The question is what should count as evidence for a safety claim before and after custody is lost. The note separates three claims that are often collapsed: what a safeguard does on the released checkpoint, what capability and refusal remain after a predefined worst-case modification, and who can independently verify the result. A proprietary alignment method may remain proprietary, including its recipe and dual-use red-team data. It should not receive regulatory credit unless the criterion is common within the crediting regime, fixed before model selection, independently testable, and independently reportable against a frozen result schema. Absence of that audit is not evidence that the safeguard is ineffective. It is a reason to withhold verified regulatory credit. After release, control does not follow the weights. It moves to deployment environment, infrastructure and access chokepoints, detection, and response. Where the environment is illegible, the remainder is residual risk, not a stop button. Authority is not causal reach. Ownership of weights is not the regulatory object. Secrecy is not evidence. This is a proposed verification and governance framework. It does not establish that any particular model or laboratory satisfies, or fails, the criteria.
Authors
- Yanush Feshter (ORCID: https://orcid.org/0009-0002-1330-7530)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-30
- DOI
- https://doi.org/10.5281/zenodo.23064519
- Primary Topic
- Regulation and Compliance Studies
- Type
- preprint