Towards Autonomous Bio-Inspired Optimization: Deep Reinforcement Learning for Adaptive Metaheuristic Orchestration
Learning-based hyper-heuristics for combinatorial optimization select from among low-level operators or tune the parameters of a single metaheuristic. However, online selection among complete population-based metaheuristics that share one evolving population and its joint configuration through a common action space have not been formulated or evaluated for binary combinatorial optimization problems. This work presents such a formulation and evaluates it on two binary domains: the multidimensional knapsack problem (MKP), and the set covering problem (SCP). A proximal policy optimization (PPO) agent orchestrates a portfolio of seven nature-inspired population-based metaheuristics, transferring the full population between methods at each decision, then acts at every fixed quantum of iterations on a fourteen-dimensional size-independent description of the search state. Over the same state, reward, and portfolio, three action spaces of increasing richness are instantiated: discrete selection, joint selection and configuration through algorithm-specific translation maps, and heterogeneous co-evolution through a simplex allocation. The three techniques are evaluated under five-fold cross-validation with twenty recorded seeds on the Chu and Beasley cb9 set (n=500 items, B=25,000 evaluations, quantum Q=25) against the best-known values of the recent literature, and the selection technique is also evaluated on 65 set-covering instances. On cb9, the selection agent is ahead of the blind round-robin and random baselines. The co-evolution agent places second, with a 0.043 percentage point gap to the best-known value behind the strongest single metaheuristic. At half the evaluation budget, the two agents rank first and second; on the Set Covering Problem, where no single method dominates, the selection agent ranks first among eleven strategies. An untrained control and a portfolio ablation show that the policy learns mainly which members to avoid; the joint selection-and-configuration agent does not yet improve on the blind baselines. The formulation, translation maps, simplex allocation, and evaluation protocol are released with this study.
Authors
- Carlos Valle (ORCID: https://orcid.org/0000-0001-7158-2069)
- Roberto Zulantay
- Giovanni Giachetti (ORCID: https://orcid.org/0000-0003-2809-5120)
- Broderick Crawford (ORCID: https://orcid.org/0000-0001-5500-0188)
- Ricardo Soto (ORCID: https://orcid.org/0000-0002-5755-6929)
- César Carrasco (ORCID: https://orcid.org/0009-0005-8179-0861)
Institutions
- Pontificia Universidad Católica de Valparaíso (CL)
- Universidad Andrés Bello (CL)
Publication Details
- Journal
- Biomimetics
- Published
- 2026-09-21
- DOI
- https://doi.org/10.3390/biomimetics11090682
- Primary Topic
- Metaheuristic Optimization Algorithms Research
- Type
- article
- Field-Weighted Citation Impact
- 0.00