Adaptive AI Inference Optimization: A Comparative Simulation Study of Static, Reactive, Forecast-Based, and Optimization-Based Autoscaling
This study presents an Adaptive AI Inference Optimizer for evaluating autoscaling strategies under changing AI inference workloads. The framework compares static provisioning, reactive autoscaling, forecast-based autoscaling, and optimization-based autoscaling using an identical synthetic workload and locally simulated inference infrastructure. The experimental evaluation uses 1,440 simulation time steps and measures simulated infrastructure cost, estimated latency, SLA violations, active instance utilization, throughput, and queue behavior. Within the specific simulation configuration used in this study, the optimization-based strategy achieved the lowest simulated total infrastructure cost of 41.10, while maintaining zero SLA violations and zero maximum queue length. Reactive autoscaling also achieved zero SLA violations but at a higher simulated cost of 60.12. Forecast-based autoscaling achieved a simulated cost of 45.65 with five SLA violations, while static provisioning recorded the highest cost of 72.00 and 195 SLA violations. All results are derived from synthetic workloads and locally simulated AI inference infrastructure. The reported cost, latency, throughput, and SLA metrics must not be interpreted as measurements from real production cloud infrastructure. The findings are therefore scoped to the workload characteristics, simulation assumptions, and strategy configurations used in this study.
Authors
- Divyashri Chinchole
Institutions
- Raisoni Group of Institutions (IN)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-04
- DOI
- https://doi.org/10.5281/zenodo.22292271
- Primary Topic
- Cloud Computing and Resource Management
- Type
- preprint