Verifying Recursive Self-Improvement: A Pre-Registered Protocol, and Results at L2 × L4
Claims of recursive self-improvement are easy to make and hard to falsify, because the system that performs the improvement is usually also the system that decides whether an improvement occurred. This work treats that as the primary problem and contributes a verification protocol for self-improvement claims, together with results obtained under it. The protocol. Six components: budget-matched controls under common random numbers; a commit-then-reveal discipline in which a candidate's identity is committed before the evidence that will judge it is disclosed; held-out evidence drawn by a secret-keyed authority and burned on use; the campaign rather than the round as the unit of analysis; calibration of the measuring instrument against a null constructed from two copies of the control; and pre-registration of every endpoint, threshold and decision rule before data collection, with analysis programs hashed at the same time. The protocol is domain-independent and re-implementable around a different loop. Results. Applied to a self-applying improvement loop across 50 independent instances, 930 governed rounds and 89,280 live deployment problems, the protocol yields a per-round advantage over the matched control of +0.0299 [+0.0189, +0.0408] (n = 40 campaigns, p = 1.2 × 10⁻⁶); five autonomy criteria holding in 299 of 300 campaign-evaluations; and, on a conjunctive test fixed in advance over 20 independent campaigns, attainment of depth L2 under autonomy L4 as scoped. The system is bit-reproducible. The central negative result is reported in full: the improvement does not compound. Across five measurements under the reference standard the promoted lineage is indistinguishable from a control lineage held to the same admission standard. Both available explanations are now eliminated — measurement noise analytically, governance asymmetry by a dedicated pre-registered study whose registration was externally timestamped before its data existed. That study also forced the retraction of the author's own prior explanation: the protocol's anti-self-deception clause, previously characterised as a tax the loop paid, is load-bearing. Removing it admitted 25% of promotions that lost to the control and made the trajectory worse. The protocol also produced four results a weaker one would have missed. It falsified a benchmark control that destroyed its own starting point. It caught a starved control arm and a leaked evidence split in the author's own runs. It forced a downward correction of the headline estimate by a third, after an apparent replication turned out to be non-independent. And it forced the retraction described above. Disclosure. The system's internal architecture is withheld. The protocol, the decision rules, the statistics, the pre-registration record with hashes, and every measured outcome — including all negative results and all defects the protocol caught in the author's own work — are disclosed in full. The supplementary archive contains the SHA-256 manifest for every pre-registration, analysis program and result document in the programme, with OpenTimestamps proofs anchoring it to the Bitcoin blockchain.
Authors
- Junai Philip Felix
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-19
- DOI
- https://doi.org/10.5281/zenodo.22845885
- Primary Topic
- Access Control and Trust
- Type
- preprint