A Cache-Aware Compression Framework: Deriving the Compression Target and Stopping Rule for Transformers on CPU-Class Hardware
In post-training compression of Transformer models, the compression ratio or pruning rate is often predetermined and does not directly take into account the memory capabilities of the target hardware. This work introduces a cache-aware post-training compression framework for CPU-class hardware. The framework derives the minimum required compression ratio from the effective weight working set of the operator or layer and the effective level-3 (L3) cache budget. The cache budget is treated as a hardware resource shared with concurrent modules rather than as a latency heuristic. If functional redundancy is present, functionally similar channels are physically removed using compensated structured pruning; in cases where final 8-bit integer (INT8) alone cannot meet the hardware target, activation-aware low-rank factorization is used. The resulting weights are quantized to per-output-channel INT8 using GPTQ, and the final configuration is evaluated by a task-level accuracy gate. In the Whisper-medium Uzbek automatic speech recognition (ASR) encoder, removing 17.1% of the feed-forward network (FFN) channels reduced the size from 300 MiB to 267 MiB, with the word error rate (WER) difference from the GPTQ-only variant (per-output-channel INT8 without structured pruning) not statistically distinguishable (ΔWER = −0.0014, 95% confidence interval (CI) [−0.0115, +0.0096]). For the full model, the cascade achieved 4.14× compression at 705 MiB with WER = 0.1833, whereas extending compression to 5.34× increased WER to 0.6101. Hardware profiling showed 2.41× fewer memory stalls and a 1.91× reduction in encoder-plus-decoder-step execution time. End-to-end response time improved by 2.36×, reducing the real audio RTF from 2.57 to 1.09. Taken together, these results support hardware-derived targeting with accuracy-guided stopping as an alternative to selecting a fixed compression ratio in advance.
Authors
- Muhammadjon Musaev (ORCID: https://orcid.org/0000-0002-9355-0954)
- Mannon Ochilov (ORCID: https://orcid.org/0000-0002-8330-0855)
- Shokhrukhmirzo Kholdorov (ORCID: https://orcid.org/0000-0003-2686-4627)
- Ilyos Khujayorov (ORCID: https://orcid.org/0000-0002-0573-6303)
- Oybek Narzullayev (ORCID: https://orcid.org/0009-0009-6654-2447)
Institutions
- Samarkand State University named after Sharof Rashidov (UZ)
- Tashkent University of Information Technology (UZ)
Publication Details
- Journal
- AI
- Published
- 2026-10-05
- DOI
- https://doi.org/10.3390/ai7100408
- Primary Topic
- Advanced Neural Network Applications
- Type
- article
- Field-Weighted Citation Impact
- 0.00