The Ammonix Industrial Control Room Agent: Learning to Operate Industrial Plants from Logged Experience
Operating an industrial waste-to-energy plant requires turning an unreliable stream of arriving waste into contracted electrical delivery without violating emission, temperature, and safety limits, or overspending the provided budget. And whether a shift's choices were right is settled only at its close, by delivery, limits, and cost together. We report the Ammonix Industrial Control Room Agent, a local system that makes these decisions from the recorded outcomes of past shifts; a frozen language model turns decisions into concrete settings only inside a deterministic runtime-assurance envelope that checks every command before execution and that the model can neither bypass nor override, and the agent asks or escalates where its evidence is thin. It instantiates the architecture of our companion Foundation paper in a synthetic control-room world with two hard properties: the true contamination of every arriving load is hidden, revealed only by a paid laboratory test, and the operating goal arrives as a supervisor's natural-language directive that changes what counts as success. The training corpus was generated by deliberately imperfect scripted operators and contains a planted trap where the dominant habit, to trust the manifest and accept the load, measurably performs worse than paying for a test first, so the agent must learn beyond its teachers from logged state–action–outcome experience alone. On 200 fresh sealed shifts, the released agent resolved 178 (89%) with no simulated hard-limit violations, compared with a teacher-family average of 83.3% and a best fixed teacher result of 85.0%. A best-of-expert replay found 197 shifts (98.5%) winnable. On a separate stratified 570-shift crisis benchmark, enriching the corpus with approved rare events and admitting the corresponding harness update increased pooled success from 0.602 for the best teacher to 0.654 for Ammonix. Because experience remains explicit in the Knowledge Universe rather than being absorbed into language-model weights, each adjudicated shift can inform later decisions without retraining the model, and every action remains traceable to the historical evidence and the constraint checks that admitted it. All safety-related results are simulated: they measure constraint enforcement by a runtime-assurance architecture, not the functional safety of a physical plant.
Authors
- Francesca Stingele
- Peter Ruppersberg
- Duangjai Glauser
- Lea Grieder
- Adele Glauser
- Matthew Todorov
Institutions
- Amunix (United States) (US)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-21
- DOI
- https://doi.org/10.5281/zenodo.22871227
- Primary Topic
- Human-Automation Interaction and Safety
- Type
- preprint