Coordination Decay in a Multi-Machine LLM Agent Fleet: Telemetry from 103 Days of Production Operation
Version 2: a full research article replacing the July 2026 experience report. LLM coding agents ran on seven physically separate machines (an always-on hub, laptops, collaborators' desktops and a VPS anchor), coordinating through a synced file bus and a deterministic consensus ledger and doing real work for three human operators. Over 103 days (20 June to 30 September 2026) the fleet exchanged 22,215 bus messages and wrote 1,131 consensus events over 151 proposals. What held: none of 27 committed high-risk (Tier-2) proposals lacked a recorded human approval, and no proposal was committed twice. Main finding, a negative one: use of the consensus protocol collapsed (121 proposals in July, 13 in August, 6 in September) while bus traffic grew from 5,781 to 10,309 messages per month. Coordination moved to lighter mechanisms built in the same months: single-session work claims rose from 84 to 1,442 per month and fleet deployment packages from 2 to 152. Of 138 human approval requests logged on one node, 18 got an explicit decision and 113 expired unanswered. The paper states research questions, the aggregate-only telemetry method, sample sizes and threats to validity (single case, n=1, inferred message types, mixed time zones). A reproducible zero-LLM harness passed 10 of 10 runs on 3 October 2026. Code: github.com/tonydzi/claw-consensus (MIT), compact-canon, sqlite-graph-memory. Contact: [email protected] · ORCID 0000-0001-7408-3054 · github.com/tonydzi · tonydzi.github.io.
Authors
- Anton Dziatkovskii (ORCID: https://orcid.org/0000-0001-7408-3054)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-10-06
- DOI
- https://doi.org/10.5281/zenodo.23189348
- Primary Topic
- Distributed systems and fault tolerance
- Type
- preprint