Level-4 Bounded Recursive Self-Improvement under a Published Definition: A Preregistered Test of a Self-Referential Improvement Mechanism
Claims of recursive self-improvement (RSI) are usually judged against criteria chosen by the people making them. We instead test a self-modifying runtime module against a published capability definition: Level 4, bounded recursive self-improvement, in the taxonomy of Liu et al. (2026). Level 4 requires a system to propose, validate and apply modifications; to rewrite its own improvement mechanism, so that the mechanism becomes self-referential; and to sustain long-term progress within a bounded domain. Two earlier stages of this programme kept the improvement mechanism fixed and therefore sit at Level 3 under that taxonomy. In the third stage, FELIX-C3, the improvement mechanism is an explicit, versioned object that contains the procedure for editing itself. The system proposes variants of the mechanism, validates each in a budget-neutral paired trial against the incumbent, and adopts it only by a margin that is itself part of the mechanism. We registered the definition, five definitional criteria, two causal criteria, the autonomy criteria, the seeds and two gates before writing any code; froze and hashed the build; and ran the confirmatory experiment once (24 previously unused seeds × 3 conditions, 32 rounds per mission, no operator). All definitional criteria passed. Every mission promoted audited revisions (24/24). Validated rewrites of the mechanism were adopted in 21 of 24 missions (60 of 166 trials), and in 17 of 24 an adopted rewrite also changed the rewriting procedure itself, although the trial that admits a rewrite cannot measure the effect of that part of it. Test solve rate rose from 27.4% to 38.0% at the midpoint and 40.1% at the end, with second-half progress of +2.1 pp (18 seeds up, 2 down, 4 unchanged; exact one-sided Wilcoxon p = 7.6 × 10⁻⁶). All 72 missions reconciled with the frozen sources and their hash chains, and all ran with zero interventions under a supervisor that passed 14 of 14 fault-scenario groups and detected 8 of 8 deliberately broken supervisor builds. The causal criteria failed, as the power gate had predicted before the run: rewriting the mechanism did not significantly outperform the same system with its mechanism frozen (+0.35 pp, p = 0.23) or with rewrites adopted at random (+0.72 pp, p = 0.079). The system therefore meets the Level-4 definition as operationalised here, but we find no evidence that its self-referential rewriting is what drives its progress.
Authors
- Junai P. Felix
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-10-04
- DOI
- https://doi.org/10.5281/zenodo.23138421
- Primary Topic
- Psychiatry, Mental Health, Neuroscience
- Type
- preprint