Adaptive Policy Learning in Situated Learning Accessibility
The first five papers of the MDAA series establish when evidence counts (scope), what a system can say about what it observed (status by proposition), what authorizes action (warrant), how provenance is traced, and how long records remain readable (standing, with aging as a declared specification). None of them lets anything move: aging is frozen before it operates and is never inferred from the trajectory it weighs. The fifth paper closes memory with an explicit refusal and names the door it does not open, the forgetting factor that adapts to prediction error, reserving for the sixth layer the case in which a declared α is adjusted with enough within-person replicates and "stops being fixed without ceasing to be declared" (Bonomo, 2026e, §6). This paper gives content to that sentence. Its question is under what conditions a system may adjust a declared policy without the policy ceasing to be declared. The regulatory regime in force for systems that learn in use answers with a declared range and a declared rule, and informs users of a change after it is made. The paper argues that this is not enough for the person the system acts upon, and proposes two gates. The first is provenance: the value in force must be reconstructible (derivable from the record by a fixed rule, with its evidence named) and legible (the derivation fits a sentence a person can read). The second, inherited from the fifth paper, is replication: an isolated disagreement between what the policy predicted and what the record showed does not say whether the world changed or the ruler was wrong, and no adjustment is admissible on it. Replication makes a disagreement admissible; it does not identify its cause. The two gates stop different things, and the paper says which. A parameter in force is then treated as a record, with three reading states (declared, adjusted, suspended) and four kinds of transition between them. The signal admitted is disagreement, measured against a norm internal to the record; performance, measured against an external norm, is justification and belongs to the next layer. Of the two kinds of disagreement only one moves the ruler: a record read as aged and confirmed on return. A record read as applicable and contradicted is a conflict, which informs and does not elect; so the layer, left to itself, gives weight back and never takes it away. What the paper claims as new is neither the ruler inside the state nor the parameter as a record, both of which have precedent, but a requirement: that legibility to the affected person be a condition of the adjustment rather than its consequence. A preregistered bench with a negative control for each rule finds the two gates separable: each control fails the gate it drops and passes the one it keeps. Of eight predictions six were met. The two that failed, one against the paper and one in an underpowered control, are reported as the preregistration classifies them, and the first gives the replication gate its price: it is paid in time.
Authors
- Hudson Augusto Rodrigues Bonomo (ORCID: https://orcid.org/0000-0003-0656-7641)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-10-08
- DOI
- https://doi.org/10.5281/zenodo.23241779
- Primary Topic
- Ethics and Social Impacts of AI
- Type
- preprint