From Execution Trajectories to Durable Cognition: Closing the Lifelong Self-Improvement Loop in Resource-Constrained Autonomous Agents

Autonomous Large Language Model (LLM) agents operating in local workstation and edge environments are fundamentally limitedby physical hardware: finite Key-Value (KV) cache memory (≤ 16GB VRAM), narrow context windows (8K–32K tokens), andgeneration compute ceilings. Under these constraints, an agent’s inability to learn from runtime execution produces an acute failuremode: Operational Amnesia. Identical compiler errors, OS shell traps, and tool misconfigurations are repeatedly encountered,investigated, and debugged across independent sessions, dissipating up to 61.4% of prompt tokens and 66.7% of conversational turnson redundant trial-and-error cycles. Existing memory architectures fail to resolve this dilemma due to multi-day batch consolidationlatency (the Amnesia Gate), query-agnostic retrieval that pollutes finite context, and unconditional retrieval reinforcement thatentrenches irrelevant knowledge in stagnation echo chambers.In this paper, we introduce the theoretical foundations and empirical realization of the IronClad Continuous LearningSubsystem, a Rust-native autonomous cognitive architecture engineered for lifelong self-improvement on consumer hardware.IronClad introduces five coordinated systems: (1) Immediate Post-Task Distillation, extracting runtime friction and recoverystrategies within ≤ 2 ms upon task termination; (2) Dynamic Task-Aware Retrieval, projecting incoming goals into a domainkeyword topology to deliver surgical context without token bloat; (3) Outcome-Grounded Reinforcement Scoring, tying memoryimportance updates to verified empirical task utility rather than retrieval frequency; (4) Dual-End Episodic Trajectory Sampling,preserving both initial intent and terminal resolutions within an O(1) token footprint; and (5) Autonomous Direct PreferenceOptimization (DPO) Trajectory Synthesis, streaming verified trajectories into ShareGPT datasets for downstream local parameteradaptation (QLoRA). We analyze the convergence properties of the Monotonic Error Suppression and Bounded ImportanceEquilibrium mechanisms under explicitly stated operational assumptions. Empirical evaluations on 20 sequential realistic compilationand refactoring tasks demonstrate that IronClad collapses recurrent error rates from 64.6% to 4.2%, cuts mean turns to resolutionby 66.7%, reduces prompt token consumption by 61.4%, and achieves a 3.4× end-to-end wall-clock speedup. Furthermore,on a strictly held-out test suite (n = 50), autonomous local QLoRA fine-tuning on IronClad-synthesized trajectories executes in12.0 minutes on consumer hardware (RTX 5060 Ti), elevating valid AST edit rates from 88.0% to 90.0% (+2.0 percentage points,representing 1 additional task out of 50 within statistical noise) while establishing the empirical scaling boundaries of parameteradaptation versus in-context procedural distillation

Authors

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-27
DOI
https://doi.org/10.5281/zenodo.23000058
Primary Topic
Multimodal Machine Learning Applications
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

From Execution Trajectories to Durable Cognition: Closing the Lifelong Self-Improvement Loop in Resource-Constrained Autonomous Agents

Wael Sahli
Zenodo (CERN European Organization for Nuclear Research)
Multimodal Machine Learning Applications
article

From Execution Trajectories to Durable Cognition: Closing the Lifelong Self-Improvement Loop in Resource-Constrained Autonomous Agents

Wael Sahli
article en

Abstract

Autonomous Large Language Model (LLM) agents operating in local workstation and edge environments are fundamentally limitedby physical hardware: finite Key-Value (KV) cache memory (≤ 16GB VRAM), narrow context windows (8K–32K tokens), andgeneration compute ceilings. Under these constraints, an agent’s inability to learn from runtime execution produces an acute failuremode: Operational Amnesia. Identical compiler errors, OS shell traps, and tool misconfigurations are repeatedly encountered,investigated, and debugged across independent sessions, dissipating up to 61.4% of prompt tokens and 66.7% of conversational turnson redundant trial-and-error cycles. Existing memory architectures fail to resolve this dilemma due to multi-day batch consolidationlatency (the Amnesia Gate), query-agnostic retrieval that pollutes finite context, and unconditional retrieval reinforcement thatentrenches irrelevant knowledge in stagnation echo chambers.In this paper, we introduce the theoretical foundations and empirical realization of the IronClad Continuous LearningSubsystem, a Rust-native autonomous cognitive architecture engineered for lifelong self-improvement on consumer hardware.IronClad introduces five coordinated systems: (1) Immediate Post-Task Distillation, extracting runtime friction and recoverystrategies within ≤ 2 ms upon task termination; (2) Dynamic Task-Aware Retrieval, projecting incoming goals into a domainkeyword topology to deliver surgical context without token bloat; (3) Outcome-Grounded Reinforcement Scoring, tying memoryimportance updates to verified empirical task utility rather than retrieval frequency; (4) Dual-End Episodic Trajectory Sampling,preserving both initial intent and terminal resolutions within an O(1) token footprint; and (5) Autonomous Direct PreferenceOptimization (DPO) Trajectory Synthesis, streaming verified trajectories into ShareGPT datasets for downstream local parameteradaptation (QLoRA). We analyze the convergence properties of the Monotonic Error Suppression and Bounded ImportanceEquilibrium mechanisms under explicitly stated operational assumptions. Empirical evaluations on 20 sequential realistic compilationand refactoring tasks demonstrate that IronClad collapses recurrent error rates from 64.6% to 4.2%, cuts mean turns to resolutionby 66.7%, reduces prompt token consumption by 61.4%, and achieves a 3.4× end-to-end wall-clock speedup. Furthermore,on a strictly held-out test suite (n = 50), autonomous local QLoRA fine-tuning on IronClad-synthesized trajectories executes in12.0 minutes on consumer hardware (RTX 5060 Ti), elevating valid AST edit rates from 88.0% to 90.0% (+2.0 percentage points,representing 1 additional task out of 50 within statistical noise) while establishing the empirical scaling boundaries of parameteradaptation versus in-context procedural distillation

Zenodo (CERN European Organization for Nuclear Research)
Openalex Percentile: Top 14%
Multimodal Machine Learning Applications
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.