nox-mem: Pain-Weighted Hybrid Memory for LLM Agents

An open-source memory layer for LLM agents, measured against five other systems on public benchmarks, in single-file SQLite stores you can host yourself. The paper reports the G3→G10d ablation trajectory, one pre-specified cross-system comparison, and the findings that cut against its own headline. What changed in v1.0.6 (2026-10-04) Scoring correction: repeated retrieved ids now count once; nine Mem0 values in Section 6 change by at most 0.004 (the largest is the LongMemEval margin, +0.119 to +0.123, as Mem0's LongMemEval nDCG@10 goes from 0.4061 to 0.4030), no ranking or sign changes. Section 6.3.2 points to the evidence dataset 10.5281/zenodo.23146656. Wording corrections in Section 6.3.2, and in the passages that repeat its conclusions (6.5, 6.7, 7.1, 8), after independent reviews. The net effect of the corpus-collision confound (e) is stated as not isolated, instead of as running against nox-mem. The task-type ablation is described as removing task types from both arms, with nox-mem's lead persisting; it no longer claims to rule out or bound the task-type contribution, since the generic run also lacked 23 of 2,370 gold chunks and logged 33 dense-search errors. The evidence dataset is described as a verifier of 132 Section 6 values plus three arithmetic differences, with twelve groups of figures it cannot recompute. No measured result changed beyond the scoring correction above. What changed in v1.0.5 (2026-10-04) Front matter only: the first page now gives the author's ORCID and contact address, this version's DOI and the concept DOI, and the code and data repository; the note on the measurement host moved from the first page to Section 5.7. No result, claim, reference or citation changed. What changed in v1.0.4 (2026-10-04) Writing pass: prose revised for plainness and concision (shorter sentences, fewer dashes and less emphasis). No number, claim, reference or citation changed; a parity script checks this mechanically over the whole manuscript, and a scoped review of the rewritten sentences restored one sentence whose scope had widened. Details in paper/CHANGELOG.md. What changed in v1.0.3 (2026-10-03; review and audit passes 2026-10-04) (A) Form and framing. A disclosure of generative AI use is added after the Conclusion, in line with arXiv's guidance: engineering and writing assistance, and read-only review by LLM-based reviewers from other model families. Product and roadmap framing is removed (go-to-market language, internal decision codes, "Autonomy pillar" labels); the pre-specified success criterion is kept and named as such. The duplicated, empty 6.8 heading is gone. (B) Corrections after review, each checked against the run artifacts and the EverMemBench paper: Zep ranks third (behind EverOS and nox-mem), not fourth; EverMemBench Table 4 has a Gemini-3-Flash column, so the 63.28% Overall is compared with MemOS on the same backbone (59.27%, +4.01 pp) and the GPT-4.1-mini comparison is labelled cross-backbone; combinations of retrieval-stage knobs are sub-additive, and the "retrieval ceiling" claims are withdrawn; the embedding-matched comparison no longer attributes the reversal to the embedder or to the architecture. (C) Metric and aggregation. F_MH is described as LLM-judged accuracy (it was called "strict EM"). nox-mem is also reported in Table 4's own aggregation (63.77%, +4.50 pp on the same backbone); in that aggregation the Gemini-2.5-flash cross-backbone margin is −0.06 pp. MemOS's F_MH is 18.88%, a GPT-4.1-mini figure. (D) Intervals and attributions. The IterB interval is IterB's own (95% CI [6.27, 9.79]; paired per-batch difference [0.25, 3.76]); the Wave C interval is recomputed with the t distribution; task setup is the leading, not established, account of the F_MH gap; HyperMem's 92.73% is a LoCoMo figure. (E) Audit pass, every statement re-checked against the code, the run artifacts and the cited papers. Section 3.4 now describes what the code does: crystallize stores caller-supplied procedures (no LLM, no promotion between chunk types), pain is fixed at ingest by a keyword rule and never raised afterwards, reflect answers on request without writing back, and nightly consolidation only extracts into topic files. EverMemBench F_MH is compared per backbone (6.02% against MemOS's 10.84% on Gemini-3-Flash). The per-category table of the embedding-matched comparison (6.4) is recomputed after a permuted LoCoMo category map was found; nox-mem still leads all five categories. EverOS outperforming nox-mem is stated in the abstract. Mem0's 66.88% on LoCoMo is an LLM-judge score, not F1, and the F1 ranking built on it is withdrawn; MuSiQue and HotpotQA reference figures are read from their source tables; several "significant" labels are corrected to what paired per-batch intervals support; the nox-mem RSS is the 399 MB measurement of 2026-05-29; smoke-run figures that no artifact reproduces are replaced by rescored ones. (F) Final review: the LoCoMo per-category retrieval cells match the archived run (single-hop 80.36%, temporal 77.96%), and an unsupported +2.8 pp date-normalization figure is replaced by the measured session-date injection (temporal F1 +15.94 pp); the MuSiQue and HotpotQA runs are described as working over each question's own candidate paragraphs, so they measure the reader, not retrieval; LightRAG's default stack is in-process storage; a Fisher test that treated paired runs as independent is withdrawn; the conclusion no longer reports the pre-specified criterion as met. (G) Tone and provenance: the comparison of Section 6 is described as pre-specified (execution plan committed to the public repository before the first run; not registered with an external registry), and the all-Gemini variant as a planned side experiment. The deployability and cost-ratio arguments of 5.7.2, 6.8 and 6.9 are cut; competitor RAM and cold-start estimates move to the supplement as the author's estimates; the observability and monitoring subsections move to the supplement. Three figures from a Mem0 re-execution whose script and per-query output were not retained are removed. The boost formula is defined once (score = base × (1 + sum of deltas)), and the per-category table of 6.4 cites the script that recomputes it. A final mechanical check restores two headings that rendered as plain text (5.2 and item F3 of 7.2), the missing EverMind-AI entry of 6.3.1, and the δ symbol of 4.1, which the previous PDF build dropped. No measured result changed in this part. The headline measurements (63.28% EverMemBench Overall, the nDCG@10 values of 6.3 and 6.3.2, KG-path 2.5 ms p50, 399 MB) are unchanged. paper/claims_check.py passes all 21 guards. The full list, with every changed number, is in paper/CHANGELOG.md. What changed in v1.0.2 (2026-09-29) Latency figures are back on archived artifacts. v1.0 quoted a 2026-06-15 re-check (KG path 2.9 ms, hybrid 653 ms p50) that was never archived. Every KG-path and standard-hybrid latency now reads from a versioned file: KG path 2.5 ms p50 (n=120) and hybrid 529 ms p50 (n=100), plus an earlier hybrid run at ~940 ms p50 with a different query mix. Ratios built on the old figure were recomputed or dropped. Two more systems measured since v1.0: EverOS and Zep, over the same corpus and the same 2,482 queries (§6.3.3, §6.3.4). Only Letta remains a documented deployment non-run. Residual confounds of the embedding-matched comparison: three declared in v1.0, four now (§6.3.2), including a retraction of the earlier account of the Mem0 version used in that run. The 5-batch protocol claim is scoped to EverMemBench (MuSiQue and HotPotQA are single full-dev runs); a LoCoMo retrieval@10 line no longer compares against Mem0's answer F1, a different metric. Related Work (§1.5) added; long appendices moved to verbatim supplements in the repository. Full list: paper/CHANGELOG.md in the repository. Abstract We introduce nox-mem, a persistent memory system for autonomous LLM agents built on one principle: pain-weighted hybrid memory with shadow discipline. Retrieval and retention are governed by an additive salience formula in which pain — an operator-assignable severity in [0.1, 1.0], otherwise fixed at ingest by a keyword rule, persisted on every chunk — is a first-class signal, and ranking changes pass a mandatory shadow phase before production activation. Each store is a single SQLite file with provider-swappable embeddings, MIT-licensed. Deployed in production since March 2026, it serves six specialized agents at KG-path p50 = 2.5 ms, $0 per KG-path query, and a 399 MB resident set in a single self-hosted process. Our central result is a pre-specified, same-corpus comparison against five competing memory systems, four of which produced head-to-head quality numbers. Under each system's native embedder nox-mem and Mem0 split — Mem0 wins LoCoMo (nDCG@10 0.469 vs 0.426), nox-mem wins LongMemEval. An embedding-matched variant (both Gemini 3072-d, n = 2,482) inverts that split in nox-mem's favour, with four residual confounds declared. EverOS, measured later over the same corpus and queries, outperforms nox-mem on both datasets (overall nDCG@10 0.646 vs 0.501), with a mandatory cross-encoder whose share of the gap is unmeasured; Zep ranks third, ahead of Mem0. On EverMemBench, nox-mem reaches 63.28% Overall with Gemini-3-flash, 4.01 pp above the 59.27% published for MemOS on the same backbone and below that backbone's 72.61% full-context basel

Authors

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-10-04
DOI
https://doi.org/10.5281/zenodo.23147633
Primary Topic
Topic Modeling
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
preprint

nox-mem: Pain-Weighted Hybrid Memory for LLM Agents

Luiz Antonio Busnello
Zenodo (CERN European Organization for Nuclear Research)
Topic Modeling
preprint

nox-mem: Pain-Weighted Hybrid Memory for LLM Agents

Luiz Antonio Busnello
preprint en

Abstract

An open-source memory layer for LLM agents, measured against five other systems on public benchmarks, in single-file SQLite stores you can host yourself. The paper reports the G3→G10d ablation trajectory, one pre-specified cross-system comparison, and the findings that cut against its own headline. What changed in v1.0.6 (2026-10-04) Scoring correction: repeated retrieved ids now count once; nine Mem0 values in Section 6 change by at most 0.004 (the largest is the LongMemEval margin, +0.119 to +0.123, as Mem0's LongMemEval nDCG@10 goes from 0.4061 to 0.4030), no ranking or sign changes. Section 6.3.2 points to the evidence dataset 10.5281/zenodo.23146656. Wording corrections in Section 6.3.2, and in the passages that repeat its conclusions (6.5, 6.7, 7.1, 8), after independent reviews. The net effect of the corpus-collision confound (e) is stated as not isolated, instead of as running against nox-mem. The task-type ablation is described as removing task types from both arms, with nox-mem's lead persisting; it no longer claims to rule out or bound the task-type contribution, since the generic run also lacked 23 of 2,370 gold chunks and logged 33 dense-search errors. The evidence dataset is described as a verifier of 132 Section 6 values plus three arithmetic differences, with twelve groups of figures it cannot recompute. No measured result changed beyond the scoring correction above. What changed in v1.0.5 (2026-10-04) Front matter only: the first page now gives the author's ORCID and contact address, this version's DOI and the concept DOI, and the code and data repository; the note on the measurement host moved from the first page to Section 5.7. No result, claim, reference or citation changed. What changed in v1.0.4 (2026-10-04) Writing pass: prose revised for plainness and concision (shorter sentences, fewer dashes and less emphasis). No number, claim, reference or citation changed; a parity script checks this mechanically over the whole manuscript, and a scoped review of the rewritten sentences restored one sentence whose scope had widened. Details in paper/CHANGELOG.md. What changed in v1.0.3 (2026-10-03; review and audit passes 2026-10-04) (A) Form and framing. A disclosure of generative AI use is added after the Conclusion, in line with arXiv's guidance: engineering and writing assistance, and read-only review by LLM-based reviewers from other model families. Product and roadmap framing is removed (go-to-market language, internal decision codes, "Autonomy pillar" labels); the pre-specified success criterion is kept and named as such. The duplicated, empty 6.8 heading is gone. (B) Corrections after review, each checked against the run artifacts and the EverMemBench paper: Zep ranks third (behind EverOS and nox-mem), not fourth; EverMemBench Table 4 has a Gemini-3-Flash column, so the 63.28% Overall is compared with MemOS on the same backbone (59.27%, +4.01 pp) and the GPT-4.1-mini comparison is labelled cross-backbone; combinations of retrieval-stage knobs are sub-additive, and the "retrieval ceiling" claims are withdrawn; the embedding-matched comparison no longer attributes the reversal to the embedder or to the architecture. (C) Metric and aggregation. F_MH is described as LLM-judged accuracy (it was called "strict EM"). nox-mem is also reported in Table 4's own aggregation (63.77%, +4.50 pp on the same backbone); in that aggregation the Gemini-2.5-flash cross-backbone margin is −0.06 pp. MemOS's F_MH is 18.88%, a GPT-4.1-mini figure. (D) Intervals and attributions. The IterB interval is IterB's own (95% CI [6.27, 9.79]; paired per-batch difference [0.25, 3.76]); the Wave C interval is recomputed with the t distribution; task setup is the leading, not established, account of the F_MH gap; HyperMem's 92.73% is a LoCoMo figure. (E) Audit pass, every statement re-checked against the code, the run artifacts and the cited papers. Section 3.4 now describes what the code does: crystallize stores caller-supplied procedures (no LLM, no promotion between chunk types), pain is fixed at ingest by a keyword rule and never raised afterwards, reflect answers on request without writing back, and nightly consolidation only extracts into topic files. EverMemBench F_MH is compared per backbone (6.02% against MemOS's 10.84% on Gemini-3-Flash). The per-category table of the embedding-matched comparison (6.4) is recomputed after a permuted LoCoMo category map was found; nox-mem still leads all five categories. EverOS outperforming nox-mem is stated in the abstract. Mem0's 66.88% on LoCoMo is an LLM-judge score, not F1, and the F1 ranking built on it is withdrawn; MuSiQue and HotpotQA reference figures are read from their source tables; several "significant" labels are corrected to what paired per-batch intervals support; the nox-mem RSS is the 399 MB measurement of 2026-05-29; smoke-run figures that no artifact reproduces are replaced by rescored ones. (F) Final review: the LoCoMo per-category retrieval cells match the archived run (single-hop 80.36%, temporal 77.96%), and an unsupported +2.8 pp date-normalization figure is replaced by the measured session-date injection (temporal F1 +15.94 pp); the MuSiQue and HotpotQA runs are described as working over each question's own candidate paragraphs, so they measure the reader, not retrieval; LightRAG's default stack is in-process storage; a Fisher test that treated paired runs as independent is withdrawn; the conclusion no longer reports the pre-specified criterion as met. (G) Tone and provenance: the comparison of Section 6 is described as pre-specified (execution plan committed to the public repository before the first run; not registered with an external registry), and the all-Gemini variant as a planned side experiment. The deployability and cost-ratio arguments of 5.7.2, 6.8 and 6.9 are cut; competitor RAM and cold-start estimates move to the supplement as the author's estimates; the observability and monitoring subsections move to the supplement. Three figures from a Mem0 re-execution whose script and per-query output were not retained are removed. The boost formula is defined once (score = base × (1 + sum of deltas)), and the per-category table of 6.4 cites the script that recomputes it. A final mechanical check restores two headings that rendered as plain text (5.2 and item F3 of 7.2), the missing EverMind-AI entry of 6.3.1, and the δ symbol of 4.1, which the previous PDF build dropped. No measured result changed in this part. The headline measurements (63.28% EverMemBench Overall, the nDCG@10 values of 6.3 and 6.3.2, KG-path 2.5 ms p50, 399 MB) are unchanged. paper/claims_check.py passes all 21 guards. The full list, with every changed number, is in paper/CHANGELOG.md. What changed in v1.0.2 (2026-09-29) Latency figures are back on archived artifacts. v1.0 quoted a 2026-06-15 re-check (KG path 2.9 ms, hybrid 653 ms p50) that was never archived. Every KG-path and standard-hybrid latency now reads from a versioned file: KG path 2.5 ms p50 (n=120) and hybrid 529 ms p50 (n=100), plus an earlier hybrid run at ~940 ms p50 with a different query mix. Ratios built on the old figure were recomputed or dropped. Two more systems measured since v1.0: EverOS and Zep, over the same corpus and the same 2,482 queries (§6.3.3, §6.3.4). Only Letta remains a documented deployment non-run. Residual confounds of the embedding-matched comparison: three declared in v1.0, four now (§6.3.2), including a retraction of the earlier account of the Mem0 version used in that run. The 5-batch protocol claim is scoped to EverMemBench (MuSiQue and HotPotQA are single full-dev runs); a LoCoMo retrieval@10 line no longer compares against Mem0's answer F1, a different metric. Related Work (§1.5) added; long appendices moved to verbatim supplements in the repository. Full list: paper/CHANGELOG.md in the repository. Abstract We introduce nox-mem, a persistent memory system for autonomous LLM agents built on one principle: pain-weighted hybrid memory with shadow discipline. Retrieval and retention are governed by an additive salience formula in which pain — an operator-assignable severity in [0.1, 1.0], otherwise fixed at ingest by a keyword rule, persisted on every chunk — is a first-class signal, and ranking changes pass a mandatory shadow phase before production activation. Each store is a single SQLite file with provider-swappable embeddings, MIT-licensed. Deployed in production since March 2026, it serves six specialized agents at KG-path p50 = 2.5 ms, $0 per KG-path query, and a 399 MB resident set in a single self-hosted process. Our central result is a pre-specified, same-corpus comparison against five competing memory systems, four of which produced head-to-head quality numbers. Under each system's native embedder nox-mem and Mem0 split — Mem0 wins LoCoMo (nDCG@10 0.469 vs 0.426), nox-mem wins LongMemEval. An embedding-matched variant (both Gemini 3072-d, n = 2,482) inverts that split in nox-mem's favour, with four residual confounds declared. EverOS, measured later over the same corpus and queries, outperforms nox-mem on both datasets (overall nDCG@10 0.646 vs 0.501), with a mandatory cross-encoder whose share of the gap is unmeasured; Zep ranks third, ahead of Mem0. On EverMemBench, nox-mem reaches 63.28% Overall with Gemini-3-flash, 4.01 pp above the 59.27% published for MemOS on the same backbone and below that backbone's 72.61% full-context basel

Zenodo (CERN European Organization for Nuclear Research)
Topic Modeling
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.