Governed Persistent Memory Under Enterprise Churn Version 1.0

Large-language-model agents face a persistent-state tradeoff: bounded context can lose information that falls outside the active window, while carrying accumulated interaction history increases prompt and reasoning burden and does not guarantee efficient use of the currently valid fact. We evaluate tokiMind, a governed persistent-memory system that treats memory as mutable operational state with explicit provenance, authority, contradiction handling, supersession, freshness, revocation, and tenant isolation. A frozen 300-hour benchmark executed 1,932 scheduled runs across three cohorts: full accumulated context, a bounded-context negative control without persistent memory, and bounded context with tokiMind. Across 350 scored recall cases per cohort, tokiMind achieved 350/350 accuracy (100%; exact 95% CI 98.95%-100.00%), full context achieved 341/350 (97.43%; 95% CI 95.17%-98.82%), and the bounded negative control achieved 0/350 (95% CI 0.00%-1.05%). At the recall level, all nine discordant full-context/tokiMind pairs favored tokiMind (exact two-sided McNemar/binomial p=0.00390625). Because the 350 recalls comprise seven repeated checkpoints on 50 entities, we also report an entity-level sensitivity analysis: tokiMind produced a perfect seven-checkpoint trajectory for 50/50 entities, while full context did so for 42/50; eight entity-level discordances favored tokiMind (exact two-sided p=0.0078125). Importantly, the nine full-context misses did not return stale or competing values: each visible answer terminated before producing the expected contact method, a pattern strongly consistent with exhaustion of the fixed 256-token generation budget. The benchmark therefore supports an efficiency-and-completion advantage under accumulated history more directly than a claim that full history selected stale truth. Relative to full context, tokiMind used 65.51% fewer input tokens and 66.60% fewer provider-reported total tokens, while incurring an 878 ms paired median end-to-end latency premium. Security testing was performed on the tokiMind cohort only: the frozen evaluator scored 84/96 checks, while post-hoc primitive evidence showed 12/12 expected-local selections and 0/12 foreign selections in the isolation cases that were mis-scored by an instrumentation/evaluator contract mismatch. All 1,932 experiment tasks completed successfully; two required one retry, and all 346 embedding tasks succeeded on their first attempt. The results support governed persistent memory as a promising state-management approach under controlled truth churn while exposing clear limitations in benchmark breadth, ceiling effects, evaluator independence, generation-budget sensitivity, and infrastructure economics.Supporting evidence from the frozen 300-hour experiment is being released as a separate Zenodo dataset record.

Authors

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-10-05
DOI
https://doi.org/10.5281/zenodo.23175184
Primary Topic
Topic Modeling
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
preprint

Governed Persistent Memory Under Enterprise Churn Version 1.0

Michael S. Miller
Zenodo (CERN European Organization for Nuclear Research)
Topic Modeling
preprint

Governed Persistent Memory Under Enterprise Churn Version 1.0

Michael S. Miller
preprint en

Abstract

Large-language-model agents face a persistent-state tradeoff: bounded context can lose information that falls outside the active window, while carrying accumulated interaction history increases prompt and reasoning burden and does not guarantee efficient use of the currently valid fact. We evaluate tokiMind, a governed persistent-memory system that treats memory as mutable operational state with explicit provenance, authority, contradiction handling, supersession, freshness, revocation, and tenant isolation. A frozen 300-hour benchmark executed 1,932 scheduled runs across three cohorts: full accumulated context, a bounded-context negative control without persistent memory, and bounded context with tokiMind. Across 350 scored recall cases per cohort, tokiMind achieved 350/350 accuracy (100%; exact 95% CI 98.95%-100.00%), full context achieved 341/350 (97.43%; 95% CI 95.17%-98.82%), and the bounded negative control achieved 0/350 (95% CI 0.00%-1.05%). At the recall level, all nine discordant full-context/tokiMind pairs favored tokiMind (exact two-sided McNemar/binomial p=0.00390625). Because the 350 recalls comprise seven repeated checkpoints on 50 entities, we also report an entity-level sensitivity analysis: tokiMind produced a perfect seven-checkpoint trajectory for 50/50 entities, while full context did so for 42/50; eight entity-level discordances favored tokiMind (exact two-sided p=0.0078125). Importantly, the nine full-context misses did not return stale or competing values: each visible answer terminated before producing the expected contact method, a pattern strongly consistent with exhaustion of the fixed 256-token generation budget. The benchmark therefore supports an efficiency-and-completion advantage under accumulated history more directly than a claim that full history selected stale truth. Relative to full context, tokiMind used 65.51% fewer input tokens and 66.60% fewer provider-reported total tokens, while incurring an 878 ms paired median end-to-end latency premium. Security testing was performed on the tokiMind cohort only: the frozen evaluator scored 84/96 checks, while post-hoc primitive evidence showed 12/12 expected-local selections and 0/12 foreign selections in the isolation cases that were mis-scored by an instrumentation/evaluator contract mismatch. All 1,932 experiment tasks completed successfully; two required one retry, and all 346 embedding tasks succeeded on their first attempt. The results support governed persistent memory as a promising state-management approach under controlled truth churn while exposing clear limitations in benchmark breadth, ceiling effects, evaluator independence, generation-budget sensitivity, and infrastructure economics.Supporting evidence from the frozen 300-hour experiment is being released as a separate Zenodo dataset record.

Zenodo (CERN European Organization for Nuclear Research)
Topic Modeling
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.