nox-mem: Pain-Weighted Hybrid Memory for LLM Agents

An open-source memory layer for LLM agents, measured against five other systems on public benchmarks, in a single-file SQLite store you can host yourself. The paper reports the G3→G10d ablation trajectory, one pre-registered cross-system comparison, and the findings that cut against its own headline. What changed in v1.0.2 (2026-09-29) Latency figures are back on archived artifacts. v1.0 quoted a 2026-06-15 re-check (KG path 2.9 ms, hybrid 653 ms p50) that was never archived. Every KG-path and standard-hybrid latency now reads from a versioned file: KG path 2.5 ms p50 (n=120) and hybrid 529 ms p50 (n=100), plus an earlier hybrid run at ~940 ms p50 with a different query mix. Ratios built on the old figure were recomputed or dropped. Two more systems measured since v1.0: EverOS and Zep, over the same corpus and the same 2,482 queries (§6.3.3, §6.3.4). Only Letta remains a documented deployment non-run. Residual confounds of the embedding-matched comparison: three declared in v1.0, four now (§6.3.2), including a retraction of the earlier account of the Mem0 version used in that run. The 5-batch protocol claim is scoped to EverMemBench (MuSiQue and HotPotQA are single full-dev runs); a LoCoMo retrieval@10 line no longer compares against Mem0's answer F1, a different metric. Related Work (§1.5) added; long appendices moved to verbatim supplements in the repository. Full list: paper/CHANGELOG.md in the repository. Abstract We introduce nox-mem, a persistent memory system for autonomous LLM agents built on one principle: pain-weighted hybrid memory with shadow discipline. Retrieval and retention are governed by an additive salience formula in which pain — an operator-assigned severity in [0.1, 1.0], persisted on every chunk — is a first-class signal, and ranking changes pass a mandatory shadow phase before production activation. The system is a single SQLite file with provider-swappable embeddings, MIT-licensed. Deployed in production since March 14, 2026, it serves six specialized agents at KG-path p50 = 2.5 ms, $0 per KG-path query, and a 399 MB resident set in a single self-hosted process. Our central result is a pre-registered, same-corpus comparison against five competing memory systems, four of which produced head-to-head quality numbers. Under each system's native embedder nox-mem and Mem0 split — Mem0 wins LoCoMo (nDCG@10 0.469 vs 0.426), nox-mem wins LongMemEval. An embedding-matched variant (both Gemini 3072-d, n = 2,482) inverts that split in nox-mem's favour, with four residual confounds declared. On EverMemBench, nox-mem reaches 63.28% Overall with Gemini-3-flash against MemOS numbers obtained on GPT-4.1-mini, so the backbones differ and this is not a state-of-the-art claim. Findings that cut against the headline Pain-weighting, the title's own signal, is not statistically significant in isolation (§7.1). It is directional. Section-aware ranking, not pain, is the dominant empirical driver (§5.1.3). On the EverMemBench F_MH multi-hop track the system sits at 3–7%, against 18.88% strict EM for the best published system on that track (§5.4). Status of this manuscript This is a preprint. It has not been peer reviewed. It was submitted to arXiv on 2026-09-03 and not accepted, with the stated reason that it "would benefit from additional review and revision that is outside of the services we provide". arXiv does not assess scientific correctness, so that sentence means the manuscript needs peer review and arXiv does not perform peer review; it is not a finding about any specific claim. An earlier version asserted state-of-the-art results on two benchmarks simultaneously. That claim was retracted on 2026-09-03–04, together with five others; the retraction list and the mechanical checks that block their reintroduction are in the repository (paper/claims_check.py). Code, data and license github.com/totobusnello/memoria-nox (MIT) holds the evaluation harness, golden sets, ablation scripts and this paper's source. The per-query result artefacts of the embedding-matched comparison are too large for the repository; they are kept off-repository, and their SHA-256 hashes are versioned in eval/q4-comparison/MANIFESTO-LASTRO.json. The PDF is reproducible from the Markdown source with ./scripts/build-paper.sh (pandoc + xelatex). This document: CC BY 4.0. A standalone build of the engine is published as the npm package nox-mem (github.com/totobusnello/nox-mem, MIT). The production instance measured in the paper runs a private, extended build of the same engine; source paths cited in the paper refer to that tree.

Authors

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-29
DOI
https://doi.org/10.5281/zenodo.23041503
Primary Topic
Multi-Agent Systems and Negotiation
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
preprint

nox-mem: Pain-Weighted Hybrid Memory for LLM Agents

Luiz Antonio Busnello
Zenodo (CERN European Organization for Nuclear Research)
Multi-Agent Systems and Negotiation
preprint

nox-mem: Pain-Weighted Hybrid Memory for LLM Agents

Luiz Antonio Busnello
preprint en

Abstract

An open-source memory layer for LLM agents, measured against five other systems on public benchmarks, in a single-file SQLite store you can host yourself. The paper reports the G3→G10d ablation trajectory, one pre-registered cross-system comparison, and the findings that cut against its own headline. What changed in v1.0.2 (2026-09-29) Latency figures are back on archived artifacts. v1.0 quoted a 2026-06-15 re-check (KG path 2.9 ms, hybrid 653 ms p50) that was never archived. Every KG-path and standard-hybrid latency now reads from a versioned file: KG path 2.5 ms p50 (n=120) and hybrid 529 ms p50 (n=100), plus an earlier hybrid run at ~940 ms p50 with a different query mix. Ratios built on the old figure were recomputed or dropped. Two more systems measured since v1.0: EverOS and Zep, over the same corpus and the same 2,482 queries (§6.3.3, §6.3.4). Only Letta remains a documented deployment non-run. Residual confounds of the embedding-matched comparison: three declared in v1.0, four now (§6.3.2), including a retraction of the earlier account of the Mem0 version used in that run. The 5-batch protocol claim is scoped to EverMemBench (MuSiQue and HotPotQA are single full-dev runs); a LoCoMo retrieval@10 line no longer compares against Mem0's answer F1, a different metric. Related Work (§1.5) added; long appendices moved to verbatim supplements in the repository. Full list: paper/CHANGELOG.md in the repository. Abstract We introduce nox-mem, a persistent memory system for autonomous LLM agents built on one principle: pain-weighted hybrid memory with shadow discipline. Retrieval and retention are governed by an additive salience formula in which pain — an operator-assigned severity in [0.1, 1.0], persisted on every chunk — is a first-class signal, and ranking changes pass a mandatory shadow phase before production activation. The system is a single SQLite file with provider-swappable embeddings, MIT-licensed. Deployed in production since March 14, 2026, it serves six specialized agents at KG-path p50 = 2.5 ms, $0 per KG-path query, and a 399 MB resident set in a single self-hosted process. Our central result is a pre-registered, same-corpus comparison against five competing memory systems, four of which produced head-to-head quality numbers. Under each system's native embedder nox-mem and Mem0 split — Mem0 wins LoCoMo (nDCG@10 0.469 vs 0.426), nox-mem wins LongMemEval. An embedding-matched variant (both Gemini 3072-d, n = 2,482) inverts that split in nox-mem's favour, with four residual confounds declared. On EverMemBench, nox-mem reaches 63.28% Overall with Gemini-3-flash against MemOS numbers obtained on GPT-4.1-mini, so the backbones differ and this is not a state-of-the-art claim. Findings that cut against the headline Pain-weighting, the title's own signal, is not statistically significant in isolation (§7.1). It is directional. Section-aware ranking, not pain, is the dominant empirical driver (§5.1.3). On the EverMemBench F_MH multi-hop track the system sits at 3–7%, against 18.88% strict EM for the best published system on that track (§5.4). Status of this manuscript This is a preprint. It has not been peer reviewed. It was submitted to arXiv on 2026-09-03 and not accepted, with the stated reason that it "would benefit from additional review and revision that is outside of the services we provide". arXiv does not assess scientific correctness, so that sentence means the manuscript needs peer review and arXiv does not perform peer review; it is not a finding about any specific claim. An earlier version asserted state-of-the-art results on two benchmarks simultaneously. That claim was retracted on 2026-09-03–04, together with five others; the retraction list and the mechanical checks that block their reintroduction are in the repository (paper/claims_check.py). Code, data and license github.com/totobusnello/memoria-nox (MIT) holds the evaluation harness, golden sets, ablation scripts and this paper's source. The per-query result artefacts of the embedding-matched comparison are too large for the repository; they are kept off-repository, and their SHA-256 hashes are versioned in eval/q4-comparison/MANIFESTO-LASTRO.json. The PDF is reproducible from the Markdown source with ./scripts/build-paper.sh (pandoc + xelatex). This document: CC BY 4.0. A standalone build of the engine is published as the npm package nox-mem (github.com/totobusnello/nox-mem, MIT). The production instance measured in the paper runs a private, extended build of the same engine; source paths cited in the paper refer to that tree.

Zenodo (CERN European Organization for Nuclear Research)
Peace, Justice and strong institutions
Multi-Agent Systems and Negotiation
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.