KiaOmni: Neighborhood-Smoothed Attention Saliency for Budgeted KV-Cache Compression
KiaOmni is a training-free KV-cache compression policy for long-context inference. It scorescontext positions from a last-layer, last-query projection-derived attention proxy, smooths that fieldwith either a Gaussian kernel (standard deviation four) or a boxcar kernel (radius eight), protectsinitial and recent positions, and retains a budgeted set of indices. The archived evaluation coversfour 7B checkpoints and seven shared-budget policies. A paired complete-case reanalysis contains5,107 input–budget configurations and 35,749 policy outcomes. At B = 512, Gaussian smoothingattains a descriptive four-checkpoint mean of 88.2% of the corresponding FullContext CORRECTrate, versus 85.7% for boxcar, 82.9% for BlockSal, 74.8% for AdaSnapKV, 71.0% for H2O, and 61.5%for SnapKV. Separate passkey and needle-in-a-haystack experiments show strong retention undertight budgets. The portable evaluation harness materializes selected indices through a fresh prefillrather than in-place cache surgery; therefore the evidence establishes the cache-selection rule andbounded retained capacity, while end-to-end online eviction performance remains a systems-validationquestion.
Authors
- Aliwey Abood
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-14
- DOI
- https://doi.org/10.5281/zenodo.22754536
- Primary Topic
- Parallel Computing and Optimization Techniques
- Type
- preprint