Towards Autonomous Bio-Inspired Optimization: Deep Reinforcement Learning for Adaptive Metaheuristic Orchestration

Learning-based hyper-heuristics for combinatorial optimization select from among low-level operators or tune the parameters of a single metaheuristic. However, online selection among complete population-based metaheuristics that share one evolving population and its joint configuration through a common action space have not been formulated or evaluated for binary combinatorial optimization problems. This work presents such a formulation and evaluates it on two binary domains: the multidimensional knapsack problem (MKP), and the set covering problem (SCP). A proximal policy optimization (PPO) agent orchestrates a portfolio of seven nature-inspired population-based metaheuristics, transferring the full population between methods at each decision, then acts at every fixed quantum of iterations on a fourteen-dimensional size-independent description of the search state. Over the same state, reward, and portfolio, three action spaces of increasing richness are instantiated: discrete selection, joint selection and configuration through algorithm-specific translation maps, and heterogeneous co-evolution through a simplex allocation. The three techniques are evaluated under five-fold cross-validation with twenty recorded seeds on the Chu and Beasley cb9 set (n=500 items, B=25,000 evaluations, quantum Q=25) against the best-known values of the recent literature, and the selection technique is also evaluated on 65 set-covering instances. On cb9, the selection agent is ahead of the blind round-robin and random baselines. The co-evolution agent places second, with a 0.043 percentage point gap to the best-known value behind the strongest single metaheuristic. At half the evaluation budget, the two agents rank first and second; on the Set Covering Problem, where no single method dominates, the selection agent ranks first among eleven strategies. An untrained control and a portfolio ablation show that the policy learns mainly which members to avoid; the joint selection-and-configuration agent does not yet improve on the blind baselines. The formulation, translation maps, simplex allocation, and evaluation protocol are released with this study.

Authors

Institutions

Publication Details

Journal
Biomimetics
Published
2026-09-21
DOI
https://doi.org/10.3390/biomimetics11090682
Primary Topic
Metaheuristic Optimization Algorithms Research
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Towards Autonomous Bio-Inspired Optimization: Deep Reinforcement Learning for Adaptive Metaheuristic Orchestration

Carlos Valle, Roberto Zulantay, Giovanni Giachetti, Broderick Crawford et al.
Biomimetics
Metaheuristic Optimization Algorithms Research
article

Towards Autonomous Bio-Inspired Optimization: Deep Reinforcement Learning for Adaptive Metaheuristic Orchestration

Carlos Valle, Roberto Zulantay, Giovanni Giachetti, Broderick Crawford, Ricardo Soto, César Carrasco
article en

Abstract

Learning-based hyper-heuristics for combinatorial optimization select from among low-level operators or tune the parameters of a single metaheuristic. However, online selection among complete population-based metaheuristics that share one evolving population and its joint configuration through a common action space have not been formulated or evaluated for binary combinatorial optimization problems. This work presents such a formulation and evaluates it on two binary domains: the multidimensional knapsack problem (MKP), and the set covering problem (SCP). A proximal policy optimization (PPO) agent orchestrates a portfolio of seven nature-inspired population-based metaheuristics, transferring the full population between methods at each decision, then acts at every fixed quantum of iterations on a fourteen-dimensional size-independent description of the search state. Over the same state, reward, and portfolio, three action spaces of increasing richness are instantiated: discrete selection, joint selection and configuration through algorithm-specific translation maps, and heterogeneous co-evolution through a simplex allocation. The three techniques are evaluated under five-fold cross-validation with twenty recorded seeds on the Chu and Beasley cb9 set (n=500 items, B=25,000 evaluations, quantum Q=25) against the best-known values of the recent literature, and the selection technique is also evaluated on 65 set-covering instances. On cb9, the selection agent is ahead of the blind round-robin and random baselines. The co-evolution agent places second, with a 0.043 percentage point gap to the best-known value behind the strongest single metaheuristic. At half the evaluation budget, the two agents rank first and second; on the Set Covering Problem, where no single method dominates, the selection agent ranks first among eleven strategies. An untrained control and a portfolio ablation show that the policy learns mainly which members to avoid; the joint selection-and-configuration agent does not yet improve on the blind baselines. The formulation, translation maps, simplex allocation, and evaluation protocol are released with this study.

BiomimeticsVol. 11(9)
Pontificia Universidad Católica de Valparaíso (CL), Universidad Andrés Bello (CL)
Peace, Justice and strong institutions
Openalex Percentile: Top 8%
Metaheuristic Optimization Algorithms Research
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.