Slice Embedding and Look-Ahead Prefetching for Memory-Constrained Large Language Model Inference

Serving large language models on consumer hardware is constrained by dynamic random-access memory (DRAM) capacity. Standard offloading frameworks swap weights sequentially across the PCI Express (PCIe) bus, incurring high transfer latency and reducing generation throughput to fractions of a token per second. Dense transformer layers cannot be skipped entirely without causing irreversible degradation of the residual stream. This paper introduces an intra-layer horizontal slicing architecture coupled with dual-purpose slice embeddings and look-ahead prefetching. Feed-forward network (FFN) layers are partitioned along the intermediate neuron dimension into independent additive sub-matrices. High-activity neurons are identified through empirical activation profiling and retained in DRAM as a persistent hot cache, while remaining neurons reside on non-volatile memory express (NVMe) solid-state storage. Each storage slice is associated with a dual slice embedding: a signed router key for look-ahead activation prediction and a low-rank linear operator that approximates residual contributions when a slice remains unloaded. During autoregressive decoding, hidden states from layer l-1 predict top-K active slices in layer l, overlapping asynchronous flash transfers with attention computation. Evaluated on Qwen2.5 architectures across 0.5B and 7B parameter scales, our system reduces NVMe read volume by 66.7% to 75.0% and cuts active DRAM requirements by 49.2% to 60.0%. On commodity consumer hardware, the engine achieves 3.05 tokens per second on Qwen2.5-0.5B, outperforming HuggingFace Accelerate disk offloading by 3.58x while maintaining 85.0% zero-shot scientific reasoning accuracy.

Authors

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-30
DOI
https://doi.org/10.5281/zenodo.23059066
Primary Topic
Parallel Computing and Optimization Techniques
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
preprint

Slice Embedding and Look-Ahead Prefetching for Memory-Constrained Large Language Model Inference

Mirza Usama Baig
Zenodo (CERN European Organization for Nuclear Research)
Parallel Computing and Optimization Techniques
preprint

Slice Embedding and Look-Ahead Prefetching for Memory-Constrained Large Language Model Inference

Mirza Usama Baig
preprint en

Abstract

Serving large language models on consumer hardware is constrained by dynamic random-access memory (DRAM) capacity. Standard offloading frameworks swap weights sequentially across the PCI Express (PCIe) bus, incurring high transfer latency and reducing generation throughput to fractions of a token per second. Dense transformer layers cannot be skipped entirely without causing irreversible degradation of the residual stream. This paper introduces an intra-layer horizontal slicing architecture coupled with dual-purpose slice embeddings and look-ahead prefetching. Feed-forward network (FFN) layers are partitioned along the intermediate neuron dimension into independent additive sub-matrices. High-activity neurons are identified through empirical activation profiling and retained in DRAM as a persistent hot cache, while remaining neurons reside on non-volatile memory express (NVMe) solid-state storage. Each storage slice is associated with a dual slice embedding: a signed router key for look-ahead activation prediction and a low-rank linear operator that approximates residual contributions when a slice remains unloaded. During autoregressive decoding, hidden states from layer l-1 predict top-K active slices in layer l, overlapping asynchronous flash transfers with attention computation. Evaluated on Qwen2.5 architectures across 0.5B and 7B parameter scales, our system reduces NVMe read volume by 66.7% to 75.0% and cuts active DRAM requirements by 49.2% to 60.0%. On commodity consumer hardware, the engine achieves 3.05 tokens per second on Qwen2.5-0.5B, outperforming HuggingFace Accelerate disk offloading by 3.58x while maintaining 85.0% zero-shot scientific reasoning accuracy.

Zenodo (CERN European Organization for Nuclear Research)
Parallel Computing and Optimization Techniques
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Slice Embedding and Look-Ahead Prefetching for Memory-Constrained Large Language Model Inference — Mirza Usama Baig · Zenodo (CERN European Organization for Nuclear Research) (2026) | TGRS Research Map | TGRS