KiaOmni: Neighborhood-Smoothed Attention Saliency for Budgeted KV-Cache Compression

KiaOmni is a training-free KV-cache compression policy for long-context inference. It scorescontext positions from a last-layer, last-query projection-derived attention proxy, smooths that fieldwith either a Gaussian kernel (standard deviation four) or a boxcar kernel (radius eight), protectsinitial and recent positions, and retains a budgeted set of indices. The archived evaluation coversfour 7B checkpoints and seven shared-budget policies. A paired complete-case reanalysis contains5,107 input–budget configurations and 35,749 policy outcomes. At B = 512, Gaussian smoothingattains a descriptive four-checkpoint mean of 88.2% of the corresponding FullContext CORRECTrate, versus 85.7% for boxcar, 82.9% for BlockSal, 74.8% for AdaSnapKV, 71.0% for H2O, and 61.5%for SnapKV. Separate passkey and needle-in-a-haystack experiments show strong retention undertight budgets. The portable evaluation harness materializes selected indices through a fresh prefillrather than in-place cache surgery; therefore the evidence establishes the cache-selection rule andbounded retained capacity, while end-to-end online eviction performance remains a systems-validationquestion.

Authors

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-14
DOI
https://doi.org/10.5281/zenodo.22754536
Primary Topic
Parallel Computing and Optimization Techniques
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
preprint

KiaOmni: Neighborhood-Smoothed Attention Saliency for Budgeted KV-Cache Compression

Aliwey Abood
Zenodo (CERN European Organization for Nuclear Research)
Parallel Computing and Optimization Techniques
preprint

KiaOmni: Neighborhood-Smoothed Attention Saliency for Budgeted KV-Cache Compression

Aliwey Abood
preprint en

Abstract

KiaOmni is a training-free KV-cache compression policy for long-context inference. It scorescontext positions from a last-layer, last-query projection-derived attention proxy, smooths that fieldwith either a Gaussian kernel (standard deviation four) or a boxcar kernel (radius eight), protectsinitial and recent positions, and retains a budgeted set of indices. The archived evaluation coversfour 7B checkpoints and seven shared-budget policies. A paired complete-case reanalysis contains5,107 input–budget configurations and 35,749 policy outcomes. At B = 512, Gaussian smoothingattains a descriptive four-checkpoint mean of 88.2% of the corresponding FullContext CORRECTrate, versus 85.7% for boxcar, 82.9% for BlockSal, 74.8% for AdaSnapKV, 71.0% for H2O, and 61.5%for SnapKV. Separate passkey and needle-in-a-haystack experiments show strong retention undertight budgets. The portable evaluation harness materializes selected indices through a fresh prefillrather than in-place cache surgery; therefore the evidence establishes the cache-selection rule andbounded retained capacity, while end-to-end online eviction performance remains a systems-validationquestion.

Zenodo (CERN European Organization for Nuclear Research)
Partnerships for the goals
Parallel Computing and Optimization Techniques
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

KiaOmni: Neighborhood-Smoothed Attention Saliency for Budgeted KV-Cache Compression — Aliwey Abood · Zenodo (CERN European Organization for Nuclear Research) (2026) | TGRS Research Map | TGRS