Stochastic Modeling and Resource Dimensioning of Multi-Cellular Edge Intelligent Systems
Edge intelligence enables the execution of AI inference tasks on computing platforms at the network edge, typically co-located with or near the radio access network rather than in centralized clouds or on mobile devices. This approach is particularly well suited for data analytics of low-latency and resource-constrained applications, where large data volumes and stringent latency constraints require tight integration of wireless access and on-site computational resources. However, the performance and cost-efficiency of such systems fundamentally depend on the joint dimensioning of wireless and computational resources prior to deployment, specially amid spatial and temporal uncertainties. Prior works largely emphasize run-time resource allocation or employ simplified network models that decouple radio access from computing infrastructure, overlooking end-to-end correlations in large-scale deployments. This paper introduces a unified stochastic framework for dimensioning multi-cellular edge-intelligent systems. We model network topology via a Poisson point process to capture randomness in user and base-station locations, incorporating inter-cell interference, distance-proportional fractional power control, and peak-power constraints. Integrating this with queueing theory and empirical profiling of AI inference workloads, we derive tractable expressions for the end-to-end offloading delay. These enable a non-convex joint optimization problem for minimizing deployment costs while enforcing statistical quality-of-service guarantees, defined not merely by averages, but by strict tail-latency and inference accuracy constraints. We prove decomposability into convex sub-problems, ensuring global optimality with zero gap. Through numerical evaluations in noise-limited and interference-limited regimes, we identify parameter regions that yield cost-efficient designs versus those that lead to severe under-utilization or unfairness across users. Key insights include the following: smaller cells reduce transmission delay but cause higher per-request computing cost due to reduced multiplexing at the servers, while larger cells exhibit the opposite trend. Moreover, network densification reduces computational costs only when frequency reuse scales with base-station density; otherwise, sparse deployments enhance fairness and efficiency in interference-limited scenarios. Overall, our analysis provides system designers with principled guidelines for scalable, QoS-aware provisioning of edge-intelligent video analytics in next-generation cellular networks.
Authors
- Jaume Anguera Peris (ORCID: https://orcid.org/0000-0003-2817-7257)
- Joakim Jaldén (ORCID: https://orcid.org/0000-0001-6630-243X)
Institutions
- KTH Royal Institute of Technology (SE)
Publication Details
- Journal
- ACM Transactions on Modeling and Performance Evaluation of Computing Systems
- Published
- 2026-10-08
- DOI
- https://doi.org/10.1145/3849089
- Primary Topic
- IoT and Edge/Fog Computing
- Type
- article
- Field-Weighted Citation Impact
- 0.00