A hierarchical reinforcement learning and predictive control framework for automated multi-variable anesthesia

BACKGROUND AND OBJECTIVE: Automated multi-variable anesthesia must balance conflicting objectives amid patient variability, safety constraints, and partial observability. The state-of-the-art robust approach uses a Genetic Algorithm (GA) to optimize Generalized Predictive Control (GPC) parameters across a population, but the resulting fixed-parameter controller is conservative and slow to adapt. Here, we develop an adaptive drug-administration framework that maintains performance while adapting to patient variability and remaining robust to surgical stimulation and clinician intervention. METHODS: A hierarchical framework is proposed in which a high-level recurrent Reinforcement Learning (RL) agent supervises a low-level multi-variable GPC by dynamically tuning its cost-function weights. Using AReS, a multi-variable anesthesia response simulator with 44 virtual patients (35 for training, 9 held out), the framework is compared with a GA-optimized GPC under surgical stimulation and anesthesiologist intervention, and its generalization is further stress-tested on 24 synthetic in-distribution and out-of-distribution patients. RESULTS: The RL-GPC framework outperformed the fixed-parameter baseline, achieving 50% shorter induction (median 228 versus 452 seconds) while keeping every monitored variable's global score below 50 for all patients, where the baseline failed in several cases. This performance held on the held-out test set and under stimuli and interventions. On the synthetic cohorts, the framework generalized robustly within the training distribution, with no patient breaching the acceptability threshold, whereas hypnotic control degraded on out-of-distribution dynamics, delineating its operational envelope. CONCLUSIONS: A hierarchical RL-GPC framework can learn a robust, adaptive policy for multi-objective automated anesthesia. These in-silico findings suggest potential benefits for patient safety, standardized care, and reduced clinician workload, and require clinical validation.

Authors

Institutions

Publication Details

Journal
Computers in Biology and Medicine
Published
2026-09-12
DOI
https://doi.org/10.1016/j.compbiomed.2026.111922
Primary Topic
Anesthesia and Sedative Agents
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

A hierarchical reinforcement learning and predictive control framework for automated multi-variable anesthesia

Francesco Trovò, Guy A. Dumont, Emiliano Tognoli, Manuela Merlo et al.
Computers in Biology and Medicine
Anesthesia and Sedative Agents
article

A hierarchical reinforcement learning and predictive control framework for automated multi-variable anesthesia

Francesco Trovò, Guy A. Dumont, Emiliano Tognoli, Manuela Merlo, Sara Hosseinirad, Alberto M. Metelli
article en

Abstract

BACKGROUND AND OBJECTIVE: Automated multi-variable anesthesia must balance conflicting objectives amid patient variability, safety constraints, and partial observability. The state-of-the-art robust approach uses a Genetic Algorithm (GA) to optimize Generalized Predictive Control (GPC) parameters across a population, but the resulting fixed-parameter controller is conservative and slow to adapt. Here, we develop an adaptive drug-administration framework that maintains performance while adapting to patient variability and remaining robust to surgical stimulation and clinician intervention. METHODS: A hierarchical framework is proposed in which a high-level recurrent Reinforcement Learning (RL) agent supervises a low-level multi-variable GPC by dynamically tuning its cost-function weights. Using AReS, a multi-variable anesthesia response simulator with 44 virtual patients (35 for training, 9 held out), the framework is compared with a GA-optimized GPC under surgical stimulation and anesthesiologist intervention, and its generalization is further stress-tested on 24 synthetic in-distribution and out-of-distribution patients. RESULTS: The RL-GPC framework outperformed the fixed-parameter baseline, achieving 50% shorter induction (median 228 versus 452 seconds) while keeping every monitored variable's global score below 50 for all patients, where the baseline failed in several cases. This performance held on the held-out test set and under stimuli and interventions. On the synthetic cohorts, the framework generalized robustly within the training distribution, with no patient breaching the acceptability threshold, whereas hypnotic control degraded on out-of-distribution dynamics, delineating its operational envelope. CONCLUSIONS: A hierarchical RL-GPC framework can learn a robust, adaptive policy for multi-objective automated anesthesia. These in-silico findings suggest potential benefits for patient safety, standardized care, and reduced clinician workload, and require clinical validation.

Computers in Biology and MedicineVol. 215
University of British Columbia (CA), Fondazione IRCCS Istituto Nazionale dei Tumori (IT), Politecnico di Milano (IT)
Openalex Percentile: Top 8%
Anesthesia and Sedative Agents
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.