The Ammonix ECG Agent: A Local Specialized AI Agent to Transparently Read and Discuss 12-Lead Electrocardiograms
Automated interpretation of the 12-lead electrocardiogram (ECG) is a high-stakes multi-label classification problem, and one that carries ambiguity the waveform alone often cannot resolve. A patient's presentation, history, and medications can each change the reading, information a deployed ECG classifier cannot ask for. Here we report the Ammonix ECG Agent, the clinical instantiation of the agent architecture introduced in our companion Foundation paper. At inference, the agent decomposes each 12-lead resting ECG into as many as 26,490 mechanistic features through its ECG tokenizer; localizes the recording in a Knowledge Universe through a classifier swarm; applies Skills (deterministic domain constraints, learned rules targeting recurrent classifier failure modes, and clarification questions); and produces a diagnostic interpretation through a frozen vision-language model that can request patient context and incorporate it into its interpretation. The tokenizer is fixed by cardiac biophysics, the classifier swarm is fitted by supervised learning on physician-documented diagnoses and frozen at deployment, and the language model is frozen. Only the Skill layer is trained, as symbolic rules in the harness rather than model weights, by Retrospective Harness Optimization with Verifiable Rewards (RHO-VR). Placing a fixed, biophysics-constrained representation before the classifier, rather than learning an opaque convolutional embedding end to end, has two independent advantages. First, the agent's explanations are auditable by design, with each claim about the recording meant to trace to a named, deterministically computed feature or a marked region of the trace. Second, because its representation is supplied rather than learned, the agent should need far fewer labeled cases. On three of the six diagnoses tested (atrial fibrillation, right bundle branch block, and Brugada syndrome), the mechanistic model leads a self-supervised convolutional neural network (CNN) on the area under the receiver operating characteristic curve (AUROC) at every training size, most widely with the fewest labeled cases. On two of the other three the CNN leads with the fewest labeled cases before the mechanistic model overtakes it. We report two experiments. On a 17-diagnosis PhysioNet task (5,108 recordings), adding an unconstrained language model to a gradient-boosted classifier lowers precision from 0.68 to 0.60 and micro-averaged F1 (micro-F1) from 0.72 to 0.69, whereas the same frozen model applying validated Skills reaches micro-F1 0.80 and 57% exact match, against 0.73 for a classifier retrained on the additional labeled data. In a scale-up to 32 diagnoses on 63,256 recordings drawn predominantly from MIMIC-IV-ECG, the classifier swarm reaches a mean AUROC of 0.950, and 0.929 against same-source negatives only. The agent runs fully offline on a single consumer laptop and supports interactive case review with the supervising clinician.
Authors
- Birgit Donner
- Francesca Stingele
- Karen Kapur
- Peter Ruppersberg
- Duangjai Glauser
- Lea Grieder
- Adele Glauser
- Matthew Todorov
- Jan Bonhoeffer
Institutions
- University Children’s Hospital Basel (CH)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-30
- DOI
- https://doi.org/10.5281/zenodo.22871231
- Primary Topic
- ECG Monitoring and Analysis
- Type
- preprint