Adaptive AI Inference Optimization: A Comparative Simulation Study of Static, Reactive, Forecast-Based, and Optimization-Based Autoscaling

This study presents an Adaptive AI Inference Optimizer for evaluating autoscaling strategies under changing AI inference workloads. The framework compares static provisioning, reactive autoscaling, forecast-based autoscaling, and optimization-based autoscaling using an identical synthetic workload and locally simulated inference infrastructure. The experimental evaluation uses 1,440 simulation time steps and measures simulated infrastructure cost, estimated latency, SLA violations, active instance utilization, throughput, and queue behavior. Within the specific simulation configuration used in this study, the optimization-based strategy achieved the lowest simulated total infrastructure cost of 41.10, while maintaining zero SLA violations and zero maximum queue length. Reactive autoscaling also achieved zero SLA violations but at a higher simulated cost of 60.12. Forecast-based autoscaling achieved a simulated cost of 45.65 with five SLA violations, while static provisioning recorded the highest cost of 72.00 and 195 SLA violations. All results are derived from synthetic workloads and locally simulated AI inference infrastructure. The reported cost, latency, throughput, and SLA metrics must not be interpreted as measurements from real production cloud infrastructure. The findings are therefore scoped to the workload characteristics, simulation assumptions, and strategy configurations used in this study.

Authors

Institutions

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-04
DOI
https://doi.org/10.5281/zenodo.22292271
Primary Topic
Cloud Computing and Resource Management
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
preprint

Adaptive AI Inference Optimization: A Comparative Simulation Study of Static, Reactive, Forecast-Based, and Optimization-Based Autoscaling

Divyashri Chinchole
Zenodo (CERN European Organization for Nuclear Research)
Cloud Computing and Resource Management
preprint

Adaptive AI Inference Optimization: A Comparative Simulation Study of Static, Reactive, Forecast-Based, and Optimization-Based Autoscaling

Divyashri Chinchole
preprint en

Abstract

This study presents an Adaptive AI Inference Optimizer for evaluating autoscaling strategies under changing AI inference workloads. The framework compares static provisioning, reactive autoscaling, forecast-based autoscaling, and optimization-based autoscaling using an identical synthetic workload and locally simulated inference infrastructure. The experimental evaluation uses 1,440 simulation time steps and measures simulated infrastructure cost, estimated latency, SLA violations, active instance utilization, throughput, and queue behavior. Within the specific simulation configuration used in this study, the optimization-based strategy achieved the lowest simulated total infrastructure cost of 41.10, while maintaining zero SLA violations and zero maximum queue length. Reactive autoscaling also achieved zero SLA violations but at a higher simulated cost of 60.12. Forecast-based autoscaling achieved a simulated cost of 45.65 with five SLA violations, while static provisioning recorded the highest cost of 72.00 and 195 SLA violations. All results are derived from synthetic workloads and locally simulated AI inference infrastructure. The reported cost, latency, throughput, and SLA metrics must not be interpreted as measurements from real production cloud infrastructure. The findings are therefore scoped to the workload characteristics, simulation assumptions, and strategy configurations used in this study.

Zenodo (CERN European Organization for Nuclear Research)
Raisoni Group of Institutions (IN)
Industry, innovation and infrastructure
Cloud Computing and Resource Management
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Adaptive AI Inference Optimization: A Comparative Simulation Study of Static, Reactive, Forecast-Based, and Optimization-Based Autoscaling — Divyashri Chinchole · Zenodo (CERN European Organization for Nuclear Research) (2026) | TGRS Research Map | TGRS