Slice Embedding and Look-Ahead Prefetching for Memory-Constrained Large Language Model Inference
Serving large language models on consumer hardware is constrained by dynamic random-access memory (DRAM) capacity. Standard offloading frameworks swap weights sequentially across the PCI Express (PCIe) bus, incurring high transfer latency and reducing generation throughput to fractions of a token per second. Dense transformer layers cannot be skipped entirely without causing irreversible degradation of the residual stream. This paper introduces an intra-layer horizontal slicing architecture coupled with dual-purpose slice embeddings and look-ahead prefetching. Feed-forward network (FFN) layers are partitioned along the intermediate neuron dimension into independent additive sub-matrices. High-activity neurons are identified through empirical activation profiling and retained in DRAM as a persistent hot cache, while remaining neurons reside on non-volatile memory express (NVMe) solid-state storage. Each storage slice is associated with a dual slice embedding: a signed router key for look-ahead activation prediction and a low-rank linear operator that approximates residual contributions when a slice remains unloaded. During autoregressive decoding, hidden states from layer l-1 predict top-K active slices in layer l, overlapping asynchronous flash transfers with attention computation. Evaluated on Qwen2.5 architectures across 0.5B and 7B parameter scales, our system reduces NVMe read volume by 66.7% to 75.0% and cuts active DRAM requirements by 49.2% to 60.0%. On commodity consumer hardware, the engine achieves 3.05 tokens per second on Qwen2.5-0.5B, outperforming HuggingFace Accelerate disk offloading by 3.58x while maintaining 85.0% zero-shot scientific reasoning accuracy.
Authors
- Mirza Usama Baig (ORCID: https://orcid.org/0009-0003-8426-0618)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-30
- DOI
- https://doi.org/10.5281/zenodo.23059066
- Primary Topic
- Parallel Computing and Optimization Techniques
- Type
- preprint