A Cache-Aware Compression Framework: Deriving the Compression Target and Stopping Rule for Transformers on CPU-Class Hardware

In post-training compression of Transformer models, the compression ratio or pruning rate is often predetermined and does not directly take into account the memory capabilities of the target hardware. This work introduces a cache-aware post-training compression framework for CPU-class hardware. The framework derives the minimum required compression ratio from the effective weight working set of the operator or layer and the effective level-3 (L3) cache budget. The cache budget is treated as a hardware resource shared with concurrent modules rather than as a latency heuristic. If functional redundancy is present, functionally similar channels are physically removed using compensated structured pruning; in cases where final 8-bit integer (INT8) alone cannot meet the hardware target, activation-aware low-rank factorization is used. The resulting weights are quantized to per-output-channel INT8 using GPTQ, and the final configuration is evaluated by a task-level accuracy gate. In the Whisper-medium Uzbek automatic speech recognition (ASR) encoder, removing 17.1% of the feed-forward network (FFN) channels reduced the size from 300 MiB to 267 MiB, with the word error rate (WER) difference from the GPTQ-only variant (per-output-channel INT8 without structured pruning) not statistically distinguishable (ΔWER = −0.0014, 95% confidence interval (CI) [−0.0115, +0.0096]). For the full model, the cascade achieved 4.14× compression at 705 MiB with WER = 0.1833, whereas extending compression to 5.34× increased WER to 0.6101. Hardware profiling showed 2.41× fewer memory stalls and a 1.91× reduction in encoder-plus-decoder-step execution time. End-to-end response time improved by 2.36×, reducing the real audio RTF from 2.57 to 1.09. Taken together, these results support hardware-derived targeting with accuracy-guided stopping as an alternative to selecting a fixed compression ratio in advance.

Authors

Institutions

Publication Details

Journal
AI
Published
2026-10-05
DOI
https://doi.org/10.3390/ai7100408
Primary Topic
Advanced Neural Network Applications
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

A Cache-Aware Compression Framework: Deriving the Compression Target and Stopping Rule for Transformers on CPU-Class Hardware

Muhammadjon Musaev, Mannon Ochilov, Shokhrukhmirzo Kholdorov, Ilyos Khujayorov et al.
AI
Advanced Neural Network Applications
article

A Cache-Aware Compression Framework: Deriving the Compression Target and Stopping Rule for Transformers on CPU-Class Hardware

Muhammadjon Musaev, Mannon Ochilov, Shokhrukhmirzo Kholdorov, Ilyos Khujayorov, Oybek Narzullayev
article en

Abstract

In post-training compression of Transformer models, the compression ratio or pruning rate is often predetermined and does not directly take into account the memory capabilities of the target hardware. This work introduces a cache-aware post-training compression framework for CPU-class hardware. The framework derives the minimum required compression ratio from the effective weight working set of the operator or layer and the effective level-3 (L3) cache budget. The cache budget is treated as a hardware resource shared with concurrent modules rather than as a latency heuristic. If functional redundancy is present, functionally similar channels are physically removed using compensated structured pruning; in cases where final 8-bit integer (INT8) alone cannot meet the hardware target, activation-aware low-rank factorization is used. The resulting weights are quantized to per-output-channel INT8 using GPTQ, and the final configuration is evaluated by a task-level accuracy gate. In the Whisper-medium Uzbek automatic speech recognition (ASR) encoder, removing 17.1% of the feed-forward network (FFN) channels reduced the size from 300 MiB to 267 MiB, with the word error rate (WER) difference from the GPTQ-only variant (per-output-channel INT8 without structured pruning) not statistically distinguishable (ΔWER = −0.0014, 95% confidence interval (CI) [−0.0115, +0.0096]). For the full model, the cascade achieved 4.14× compression at 705 MiB with WER = 0.1833, whereas extending compression to 5.34× increased WER to 0.6101. Hardware profiling showed 2.41× fewer memory stalls and a 1.91× reduction in encoder-plus-decoder-step execution time. End-to-end response time improved by 2.36×, reducing the real audio RTF from 2.57 to 1.09. Taken together, these results support hardware-derived targeting with accuracy-guided stopping as an alternative to selecting a fixed compression ratio in advance.

AIVol. 7(10)
Samarkand State University named after Sharof Rashidov (UZ), Tashkent University of Information Technology (UZ)
Openalex Percentile: Top 14%
Advanced Neural Network Applications
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.