How Much Oracle Is Needed in LLM-Driven Evolution? Separating Selection and Mutation Feedback under Surrogate Evaluation

LLM-driven evolutionary search has produced strong algorithms and scientific results, but its successes typically rely on an evaluator that is well aligned with the true objective. When evaluation is expensive or subjective, a cheaper surrogate may replace part of this oracle access. In conventional evolutionary computation, surrogate error primarily changes selection. In LLM-driven evolution, evaluator outputs may also be included in mutation prompts, creating a second path through which surrogate error can alter search. We separate these paths with two controls: the fraction of oracle verification used for selection, λ_sel, and the probability that a mutation prompt receives oracle rather than surrogate feedback, λ_fb. We evaluate a FunSearch-style system using gpt-oss-20b on a synthetic hierarchical policy task with a known oracle, a deliberately misspecified surrogate, and an unseen held-out probe set. Across three runs per condition, all-oracle evolution reaches 63.5 ± 1.0% of the held-out oracle ceiling. Replacing only mutation feedback with the surrogate reduces this to 49.5 ± 2.8%; replacing only selection reduces it to 34.5 ± 3.0%; replacing both reduces it to 27.6 ± 0.8%. When the two fractions are coupled (λ_sel = λ_fb), a value of one half recovers 60.7 ± 3.3%, close to the all-oracle endpoint. These results show that selection and mutation feedback are distinct channels, with selection dominant in this setting and oracle feedback providing an additional benefit. The study is intentionally diagnostic: it covers one model, one synthetic task, and three runs per condition, and does not establish a universal oracle threshold.

Authors

Institutions

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-05
DOI
https://doi.org/10.5281/zenodo.22325776
Primary Topic
Evolutionary Algorithms and Applications
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

How Much Oracle Is Needed in LLM-Driven Evolution? Separating Selection and Mutation Feedback under Surrogate Evaluation

Ryosuke Takata, Tomoya Hirayama
Zenodo (CERN European Organization for Nuclear Research)
Evolutionary Algorithms and Applications
article

How Much Oracle Is Needed in LLM-Driven Evolution? Separating Selection and Mutation Feedback under Surrogate Evaluation

Ryosuke Takata, Tomoya Hirayama
article en

Abstract

LLM-driven evolutionary search has produced strong algorithms and scientific results, but its successes typically rely on an evaluator that is well aligned with the true objective. When evaluation is expensive or subjective, a cheaper surrogate may replace part of this oracle access. In conventional evolutionary computation, surrogate error primarily changes selection. In LLM-driven evolution, evaluator outputs may also be included in mutation prompts, creating a second path through which surrogate error can alter search. We separate these paths with two controls: the fraction of oracle verification used for selection, λ_sel, and the probability that a mutation prompt receives oracle rather than surrogate feedback, λ_fb. We evaluate a FunSearch-style system using gpt-oss-20b on a synthetic hierarchical policy task with a known oracle, a deliberately misspecified surrogate, and an unseen held-out probe set. Across three runs per condition, all-oracle evolution reaches 63.5 ± 1.0% of the held-out oracle ceiling. Replacing only mutation feedback with the surrogate reduces this to 49.5 ± 2.8%; replacing only selection reduces it to 34.5 ± 3.0%; replacing both reduces it to 27.6 ± 0.8%. When the two fractions are coupled (λ_sel = λ_fb), a value of one half recovers 60.7 ± 3.3%, close to the all-oracle endpoint. These results show that selection and mutation feedback are distinct channels, with selection dominant in this setting and oracle feedback providing an additional benefit. The study is intentionally diagnostic: it covers one model, one synthetic task, and three runs per condition, and does not establish a universal oracle threshold.

Zenodo (CERN European Organization for Nuclear Research)
Tsuchiura City Museum (JP), The University of Tokyo (JP)
Openalex Percentile: Top 8%
Evolutionary Algorithms and Applications
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

How Much Oracle Is Needed in LLM-Driven Evolution? Separating Selection and Mutation Feedback under Surrogate Evaluation — Ryosuke Takata, Tomoya Hirayama · Zenodo (CERN European Organization for Nuclear Research) (2026) | TGRS Research Map | TGRS