Prompt-Shape- and Prefix-Cache-Aware CPU/NPU Routing for Large Language Model Inference on Edge SoCs
Large language model (LLM) inference on edge systems must exploit heterogeneous central processing unit (CPU) and neural processing unit (NPU) resources under tight memory and latency constraints. Because prefilling and autoregressive decoding have different computational characteristics, the fastest backend depends on the request shape and prefix-cache reuse. This work proposes request-level CPU/NPU routing for the RK3588 system on chip (SoC). Prompt-shape-aware routing (PSR) uses prompt length and generation budget to select a backend for no-cache requests, while prefix-cache-aware routing (PCAR) additionally models reusable-prefix length, remaining prefill workload, full-context effects, and cache-path state with backend-specific latency models. On 49 no-cache held-out workloads, PSR achieves 97.96% backend-selection accuracy and a relative oracle loss of 0.121% ± 0.010%. On 45 independent cache-aware held-out workloads, accounting for the remaining prefill workload reduces relative oracle loss from 4.926% ± 0.378% for PSR(raw) to 0.425% ± 0.060% for PSR(eff). The complete PCAR model further reduces it to 0.205% ± 0.043%, improves selection accuracy from 93.33% to 95.56%, and provides an additional 0.22% cumulative-latency reduction over that of PSR(eff). These results show that jointly modeling request shape and prefix-cache state enables near-oracle heterogeneous LLM inference routing on edge SoCs.
Authors
- Kouan Hao
- Liang Wang (ORCID: https://orcid.org/0000-0002-0061-5502)
- Yitian Mao
- Qiuchang Han
- Qinyu Guo
Institutions
- Shanghai Academy of Spaceflight Technology (CN)
Publication Details
- Journal
- Electronics
- Published
- 2026-10-09
- DOI
- https://doi.org/10.3390/electronics15204600
- Primary Topic
- Parallel Computing and Optimization Techniques
- Type
- article
- Field-Weighted Citation Impact
- 0.00