Auditing Architecture Search and Accepted Edits in Self-Modifying Systems

Accepting an edit that does not reduce a validation score does not establish that the edit improves a self-modifying system. We examine this distinction using a discrete skill graph and locally deployed language models. On the skill graph, supplying the reference architecture and optimizing its parameters yields 15.6 of 16 tasks solved, compared with 9.2 for joint architecture–parameter search and 4.2 for a shuffled-fitness control. In a paired scale scan with a fixed expected flip budget, recovery of a deleted critical edge together with full task performance falls from 12/12 runs at 35 candidate edges to 2/12 at 2,331. This observation is specific to the tested representation and budget; it does not distinguish a new selection mechanism from reduced mutation accessibility. In three-arm language-model experiments with internally frozen decision rules, an expert-initialized arm exceeds a programmatic random-edit baseline by 4.9 of 20 test tasks at Qwen3-14B (95% CI [2.76, 7.04]). At Qwen3-30B-A3B, a post-hoc format-and-token-budget repair eliminates directed-arm parse failures, but neither an advantage nor equivalence is established (+0.2 tasks, 95% CI [−0.96, 1.36]). All 100 random edits in the 14B baseline are accepted with zero observed validation gain. The experiments support separating acceptance counts, measured score changes, and held-out performance. They also identify limits arising from unequal prior information, unmatched text-edit operators, self-report classification, and evaluation variability. Code, event records, and supplementary reanalyses accompany the study.

Authors

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-10-05
DOI
https://doi.org/10.5281/zenodo.23158693
Primary Topic
Evolutionary Algorithms and Applications
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

Auditing Architecture Search and Accepted Edits in Self-Modifying Systems

Yulong Li
Zenodo (CERN European Organization for Nuclear Research)
Evolutionary Algorithms and Applications
article

Auditing Architecture Search and Accepted Edits in Self-Modifying Systems

Yulong Li
article en

Abstract

Accepting an edit that does not reduce a validation score does not establish that the edit improves a self-modifying system. We examine this distinction using a discrete skill graph and locally deployed language models. On the skill graph, supplying the reference architecture and optimizing its parameters yields 15.6 of 16 tasks solved, compared with 9.2 for joint architecture–parameter search and 4.2 for a shuffled-fitness control. In a paired scale scan with a fixed expected flip budget, recovery of a deleted critical edge together with full task performance falls from 12/12 runs at 35 candidate edges to 2/12 at 2,331. This observation is specific to the tested representation and budget; it does not distinguish a new selection mechanism from reduced mutation accessibility. In three-arm language-model experiments with internally frozen decision rules, an expert-initialized arm exceeds a programmatic random-edit baseline by 4.9 of 20 test tasks at Qwen3-14B (95% CI [2.76, 7.04]). At Qwen3-30B-A3B, a post-hoc format-and-token-budget repair eliminates directed-arm parse failures, but neither an advantage nor equivalence is established (+0.2 tasks, 95% CI [−0.96, 1.36]). All 100 random edits in the 14B baseline are accepted with zero observed validation gain. The experiments support separating acceptance counts, measured score changes, and held-out performance. They also identify limits arising from unequal prior information, unmatched text-edit operators, self-report classification, and evaluation variability. Code, event records, and supplementary reanalyses accompany the study.

Zenodo (CERN European Organization for Nuclear Research)
Openalex Percentile: Top 10%
Evolutionary Algorithms and Applications
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Auditing Architecture Search and Accepted Edits in Self-Modifying Systems — Yulong Li · Zenodo (CERN European Organization for Nuclear Research) (2026) | TGRS Research Map | TGRS