Continuous Implicit Manifold Attention (CIMA)
Transformer-based Large Language Models (LLMs) suffer from memory bottlenecks during long-context inference due to the $O(N)$ spatial footprint of the Key-Value (KV) cache. We propose Continuous Implicit Manifold Attention (CIMA), a novel caching architecture that approximates discrete sequence keys as continuous implicit neural fields using Sinusoidal Representation Networks (SIRENs). To prevent information loss at high-entropy token boundaries, CIMA incorporates a Sparse Residual Anchor Map ($\mathcal{A}$) that selectively retains exact key representations for the top 2% highest reconstruction error states. Evaluated across sequence lengths up to 32,000 tokens, CIMA achieves up to 97.4% VRAM footprint reduction compared to standard FP16 KV caching while maintaining 100% retrieval accuracy on long-context Needle-In-A-Haystack (NIAH) benchmarks.
Authors
- Sharvesh Sathish Kumar
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-30
- DOI
- https://doi.org/10.5281/zenodo.23049831
- Primary Topic
- Topic Modeling
- Type
- preprint