Delay-aware adaptive chemotherapy for tumor-immune dynamics using proximal policy optimization

The design of adaptive chemotherapy is complicated by nonlinear tumor–immune interactions, delayed drug transport and uncertainty in state evolution. This study develops a delay-aware computational framework that couples a normalized normal–tumor–immune model with central and peripheral pharmacokinetic compartments. A proximal policy optimization (PPO) agent observes the population states, their rates of change and both drug concentrations, and selects a bounded continuous infusion action. Pharmacodynamic killing is driven by peripheral exposure, thereby distinguishing systemic administration from the concentration acting at the tumor site. Positivity and boundedness of the deterministic biological–pharmacokinetic system are established for non-negative initial states and bounded infusion. The learning objective jointly penalizes tumor burden, positive tumor growth, cumulative exposure, abrupt normalized-action variation and depletion of normal and immune populations. Numerical experiments compare untreated and periodic regimens, one- and two-compartment environments, nominal and stochastic transitions, and PPO with two actor–critic baselines. Explicit distribution dynamics change the learned policy from a sustained high-action plateau to an early loading phase followed by gradual tapering. Moderate transition randomization yields the smallest reported tumor-tracking errors and reduced trajectory dispersion, whereas stronger perturbations substantially widen the uncertainty bands. These results identify pharmacokinetic representation as a consequential component of learning-based treatment design. The framework is an in-silico mechanism study rather than a clinical dosing recommendation and requires drug-specific calibration, hard toxicity constraints, complete statistical evaluation and biological validation.

Authors

Publication Details

Journal
International Journal of Modern Physics C
Published
2026-09-04
DOI
https://doi.org/10.1142/s0129183127501518
Primary Topic
Mathematical Biology Tumor Growth
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Delay-aware adaptive chemotherapy for tumor-immune dynamics using proximal policy optimization

Ruru Ma, Shutao Jiang, Mingliu Zhu
International Journal of Modern Physics C
Mathematical Biology Tumor Growth
article

Delay-aware adaptive chemotherapy for tumor-immune dynamics using proximal policy optimization

Ruru Ma, Shutao Jiang, Mingliu Zhu
article en

Abstract

The design of adaptive chemotherapy is complicated by nonlinear tumor–immune interactions, delayed drug transport and uncertainty in state evolution. This study develops a delay-aware computational framework that couples a normalized normal–tumor–immune model with central and peripheral pharmacokinetic compartments. A proximal policy optimization (PPO) agent observes the population states, their rates of change and both drug concentrations, and selects a bounded continuous infusion action. Pharmacodynamic killing is driven by peripheral exposure, thereby distinguishing systemic administration from the concentration acting at the tumor site. Positivity and boundedness of the deterministic biological–pharmacokinetic system are established for non-negative initial states and bounded infusion. The learning objective jointly penalizes tumor burden, positive tumor growth, cumulative exposure, abrupt normalized-action variation and depletion of normal and immune populations. Numerical experiments compare untreated and periodic regimens, one- and two-compartment environments, nominal and stochastic transitions, and PPO with two actor–critic baselines. Explicit distribution dynamics change the learned policy from a sustained high-action plateau to an early loading phase followed by gradual tapering. Moderate transition randomization yields the smallest reported tumor-tracking errors and reduced trajectory dispersion, whereas stronger perturbations substantially widen the uncertainty bands. These results identify pharmacokinetic representation as a consequential component of learning-based treatment design. The framework is an in-silico mechanism study rather than a clinical dosing recommendation and requires drug-specific calibration, hard toxicity constraints, complete statistical evaluation and biological validation.

International Journal of Modern Physics C
Openalex Percentile: Top 11%
Mathematical Biology Tumor Growth
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Delay-aware adaptive chemotherapy for tumor-immune dynamics using proximal policy optimization — Ruru Ma, Shutao Jiang, et al. · International Journal of Modern Physics C (2026) | TGRS Research Map | TGRS