Prompt-Shape- and Prefix-Cache-Aware CPU/NPU Routing for Large Language Model Inference on Edge SoCs

Large language model (LLM) inference on edge systems must exploit heterogeneous central processing unit (CPU) and neural processing unit (NPU) resources under tight memory and latency constraints. Because prefilling and autoregressive decoding have different computational characteristics, the fastest backend depends on the request shape and prefix-cache reuse. This work proposes request-level CPU/NPU routing for the RK3588 system on chip (SoC). Prompt-shape-aware routing (PSR) uses prompt length and generation budget to select a backend for no-cache requests, while prefix-cache-aware routing (PCAR) additionally models reusable-prefix length, remaining prefill workload, full-context effects, and cache-path state with backend-specific latency models. On 49 no-cache held-out workloads, PSR achieves 97.96% backend-selection accuracy and a relative oracle loss of 0.121% ± 0.010%. On 45 independent cache-aware held-out workloads, accounting for the remaining prefill workload reduces relative oracle loss from 4.926% ± 0.378% for PSR(raw) to 0.425% ± 0.060% for PSR(eff). The complete PCAR model further reduces it to 0.205% ± 0.043%, improves selection accuracy from 93.33% to 95.56%, and provides an additional 0.22% cumulative-latency reduction over that of PSR(eff). These results show that jointly modeling request shape and prefix-cache state enables near-oracle heterogeneous LLM inference routing on edge SoCs.

Authors

Institutions

Publication Details

Journal
Electronics
Published
2026-10-09
DOI
https://doi.org/10.3390/electronics15204600
Primary Topic
Parallel Computing and Optimization Techniques
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

Prompt-Shape- and Prefix-Cache-Aware CPU/NPU Routing for Large Language Model Inference on Edge SoCs

Kouan Hao, Liang Wang, Yitian Mao, Qiuchang Han et al.
Electronics
Parallel Computing and Optimization Techniques
article

Prompt-Shape- and Prefix-Cache-Aware CPU/NPU Routing for Large Language Model Inference on Edge SoCs

Kouan Hao, Liang Wang, Yitian Mao, Qiuchang Han, Qinyu Guo
article en

Abstract

Large language model (LLM) inference on edge systems must exploit heterogeneous central processing unit (CPU) and neural processing unit (NPU) resources under tight memory and latency constraints. Because prefilling and autoregressive decoding have different computational characteristics, the fastest backend depends on the request shape and prefix-cache reuse. This work proposes request-level CPU/NPU routing for the RK3588 system on chip (SoC). Prompt-shape-aware routing (PSR) uses prompt length and generation budget to select a backend for no-cache requests, while prefix-cache-aware routing (PCAR) additionally models reusable-prefix length, remaining prefill workload, full-context effects, and cache-path state with backend-specific latency models. On 49 no-cache held-out workloads, PSR achieves 97.96% backend-selection accuracy and a relative oracle loss of 0.121% ± 0.010%. On 45 independent cache-aware held-out workloads, accounting for the remaining prefill workload reduces relative oracle loss from 4.926% ± 0.378% for PSR(raw) to 0.425% ± 0.060% for PSR(eff). The complete PCAR model further reduces it to 0.205% ± 0.043%, improves selection accuracy from 93.33% to 95.56%, and provides an additional 0.22% cumulative-latency reduction over that of PSR(eff). These results show that jointly modeling request shape and prefix-cache state enables near-oracle heterogeneous LLM inference routing on edge SoCs.

ElectronicsVol. 15(20)
Shanghai Academy of Spaceflight Technology (CN)
Openalex Percentile: Top 8%
Parallel Computing and Optimization Techniques
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Prompt-Shape- and Prefix-Cache-Aware CPU/NPU Routing for Large Language Model Inference on Edge SoCs — Kouan Hao, Liang Wang, et al. · Electronics (2026) | TGRS Research Map | TGRS