Beyond the Model: Runesmith, an Open Runtime That Improves Software and Itself, and the Instrument–Substrate Hypothesis

Agent systems rent competence from a model one call at a time. We present Runesmith, an open-source runtime (Apache-2.0) in which models propose, a small deterministic kernel admits, and the runtime changes its own code only on evidence. In one iteration of its Kaizen loop, Runesmith chose a weakness from its own telemetry and a model rewrote its repair organ, the code it repairs software with. With the same cheap repair model and budget, the new organ repaired more than the organ it replaced, 57 against 35 of 162 sessions on 54 fresh tasks, in about half the time per repair (SR7; exact p = 0.000845, below its preregistered Bonferroni bar of 0.0036 and the bar for all 23 sealed comparisons run). In a further sealed test on three large public repositories with one free model (LOC1), the self-improved organ repaired more regressions than the organ Runesmith shipped with (42 against 17 of 168 sessions, p = 0.00278, significant at its Bonferroni level for its three tests); a hand-written trace-aware localization rule repaired more still (65 of 168, mirrored p = 0.00077), so on this family the learned gain is one that a careful designer can also write by hand. In a post-hoc count, 123 of the 124 repairs came with the faulty file in the model's view: what Runesmith kept is a localization policy. LOC1's protocol and result wording were sealed before any outcome existed and publicly timestamped hours before the analysis. Runesmith changed its own problem-solving code and the change helped under seal, while the improvement process itself stayed unchanged; we call this, and only this, a nudge towards recursive self-improvement. With no model available in 104 sealed scenarios, it made no unauthorized change (15 effects lacked their own ledger event, added in 1.0.0). Six of seven earlier sealed tests did not show their effect; all are reported, with code, data and recomputations.

Authors

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-10-06
DOI
https://doi.org/10.5281/zenodo.23196762
Primary Topic
Software Engineering Research
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
preprint

Beyond the Model: Runesmith, an Open Runtime That Improves Software and Itself, and the Instrument–Substrate Hypothesis

Lars O. Horpestad
Zenodo (CERN European Organization for Nuclear Research)
Software Engineering Research
preprint

Beyond the Model: Runesmith, an Open Runtime That Improves Software and Itself, and the Instrument–Substrate Hypothesis

Lars O. Horpestad
preprint en

Abstract

Agent systems rent competence from a model one call at a time. We present Runesmith, an open-source runtime (Apache-2.0) in which models propose, a small deterministic kernel admits, and the runtime changes its own code only on evidence. In one iteration of its Kaizen loop, Runesmith chose a weakness from its own telemetry and a model rewrote its repair organ, the code it repairs software with. With the same cheap repair model and budget, the new organ repaired more than the organ it replaced, 57 against 35 of 162 sessions on 54 fresh tasks, in about half the time per repair (SR7; exact p = 0.000845, below its preregistered Bonferroni bar of 0.0036 and the bar for all 23 sealed comparisons run). In a further sealed test on three large public repositories with one free model (LOC1), the self-improved organ repaired more regressions than the organ Runesmith shipped with (42 against 17 of 168 sessions, p = 0.00278, significant at its Bonferroni level for its three tests); a hand-written trace-aware localization rule repaired more still (65 of 168, mirrored p = 0.00077), so on this family the learned gain is one that a careful designer can also write by hand. In a post-hoc count, 123 of the 124 repairs came with the faulty file in the model's view: what Runesmith kept is a localization policy. LOC1's protocol and result wording were sealed before any outcome existed and publicly timestamped hours before the analysis. Runesmith changed its own problem-solving code and the change helped under seal, while the improvement process itself stayed unchanged; we call this, and only this, a nudge towards recursive self-improvement. With no model available in 104 sealed scenarios, it made no unauthorized change (15 effects lacked their own ledger event, added in 1.0.0). Six of seven earlier sealed tests did not show their effect; all are reported, with code, data and recomputations.

Zenodo (CERN European Organization for Nuclear Research)
Software Engineering Research
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.