Shared Native Memory: Expanding Knowledge Without Retraining
Abstract—Aimee provides model-agnostic native memory:shared knowledge consumed through attention tensors, withfixed model weights and no memory text in the task prompt.A completed 47,065-response comparison matches text retrievalat 100% on 10,000 general and 625 personal questions. Bothconditions retain 100% accuracy with up to 64 selected records.At 64 general records, native memory removes a median 4,580prompt tokens and reduces median time to first token from2,565 to 156 ms, using prepared host-memory state. Earlierstudies establish live corrections and knowledge transfer froma 27-billion-parameter donor to a 12-billion-parameter receiver.Three independent retained-learning replications improve from apooled 955 to 1,128 correct answers out of 1,440, without weightupdates. We have also verified compatibility with DeepSeek-V4’sstructurally distinct compressed attention, separately from thescored benchmarks. An installed plugin now serves native memory through unmodified vLLM on AMD and NVIDIA hardware.Six selected repository-repair cases yield two successes withoutmemory, three with text memory and four with native memory;an unchanged-memory repeat yields five. These small deploymentstudies establish integration behavior, not a benchmark-wideeffect. The paper reports knowledge consumption and retainedlearning, not changed reasoning weights. Protocols, preparationcosts, unsuccessful attempts and storage limits accompany theresults.
Authors
- Jared Bailes
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-10-03
- DOI
- https://doi.org/10.5281/zenodo.23120104
- Primary Topic
- Topic Modeling
- Type
- preprint