The Ammonix ECG Agent: A Local Specialized AI Agent to Transparently Read and Discuss 12-Lead Electrocardiograms

Automated interpretation of the 12-lead electrocardiogram (ECG) is a high-stakes multi-label classification problem, and one that carries ambiguity the waveform alone often cannot resolve. A patient's presentation, history, and medications can each change the reading, information a deployed ECG classifier cannot ask for. Here we report the Ammonix ECG Agent, the clinical instantiation of the agent architecture introduced in our companion Foundation paper. At inference, the agent decomposes each 12-lead resting ECG into as many as 26,490 mechanistic features through its ECG tokenizer; localizes the recording in a Knowledge Universe through a classifier swarm; applies Skills (deterministic domain constraints, learned rules targeting recurrent classifier failure modes, and clarification questions); and produces a diagnostic interpretation through a frozen vision-language model that can request patient context and incorporate it into its interpretation. The tokenizer is fixed by cardiac biophysics, the classifier swarm is fitted by supervised learning on physician-documented diagnoses and frozen at deployment, and the language model is frozen. Only the Skill layer is trained, as symbolic rules in the harness rather than model weights, by Retrospective Harness Optimization with Verifiable Rewards (RHO-VR). Placing a fixed, biophysics-constrained representation before the classifier, rather than learning an opaque convolutional embedding end to end, has two independent advantages. First, the agent's explanations are auditable by design, with each claim about the recording meant to trace to a named, deterministically computed feature or a marked region of the trace. Second, because its representation is supplied rather than learned, the agent should need far fewer labeled cases. On three of the six diagnoses tested (atrial fibrillation, right bundle branch block, and Brugada syndrome), the mechanistic model leads a self-supervised convolutional neural network (CNN) on the area under the receiver operating characteristic curve (AUROC) at every training size, most widely with the fewest labeled cases. On two of the other three the CNN leads with the fewest labeled cases before the mechanistic model overtakes it. We report two experiments. On a 17-diagnosis PhysioNet task (5,108 recordings), adding an unconstrained language model to a gradient-boosted classifier lowers precision from 0.68 to 0.60 and micro-averaged F1 (micro-F1) from 0.72 to 0.69, whereas the same frozen model applying validated Skills reaches micro-F1 0.80 and 57% exact match, against 0.73 for a classifier retrained on the additional labeled data. In a scale-up to 32 diagnoses on 63,256 recordings drawn predominantly from MIMIC-IV-ECG, the classifier swarm reaches a mean AUROC of 0.950, and 0.929 against same-source negatives only. The agent runs fully offline on a single consumer laptop and supports interactive case review with the supervising clinician.

Authors

Institutions

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-30
DOI
https://doi.org/10.5281/zenodo.22871231
Primary Topic
ECG Monitoring and Analysis
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
preprint

The Ammonix ECG Agent: A Local Specialized AI Agent to Transparently Read and Discuss 12-Lead Electrocardiograms

Birgit Donner, Francesca Stingele, Karen Kapur, Peter Ruppersberg et al.
Zenodo (CERN European Organization for Nuclear Research)
ECG Monitoring and Analysis
preprint

The Ammonix ECG Agent: A Local Specialized AI Agent to Transparently Read and Discuss 12-Lead Electrocardiograms

Birgit Donner, Francesca Stingele, Karen Kapur, Peter Ruppersberg, Duangjai Glauser, Lea Grieder, Adele Glauser, Matthew Todorov, Jan Bonhoeffer
preprint en

Abstract

Automated interpretation of the 12-lead electrocardiogram (ECG) is a high-stakes multi-label classification problem, and one that carries ambiguity the waveform alone often cannot resolve. A patient's presentation, history, and medications can each change the reading, information a deployed ECG classifier cannot ask for. Here we report the Ammonix ECG Agent, the clinical instantiation of the agent architecture introduced in our companion Foundation paper. At inference, the agent decomposes each 12-lead resting ECG into as many as 26,490 mechanistic features through its ECG tokenizer; localizes the recording in a Knowledge Universe through a classifier swarm; applies Skills (deterministic domain constraints, learned rules targeting recurrent classifier failure modes, and clarification questions); and produces a diagnostic interpretation through a frozen vision-language model that can request patient context and incorporate it into its interpretation. The tokenizer is fixed by cardiac biophysics, the classifier swarm is fitted by supervised learning on physician-documented diagnoses and frozen at deployment, and the language model is frozen. Only the Skill layer is trained, as symbolic rules in the harness rather than model weights, by Retrospective Harness Optimization with Verifiable Rewards (RHO-VR). Placing a fixed, biophysics-constrained representation before the classifier, rather than learning an opaque convolutional embedding end to end, has two independent advantages. First, the agent's explanations are auditable by design, with each claim about the recording meant to trace to a named, deterministically computed feature or a marked region of the trace. Second, because its representation is supplied rather than learned, the agent should need far fewer labeled cases. On three of the six diagnoses tested (atrial fibrillation, right bundle branch block, and Brugada syndrome), the mechanistic model leads a self-supervised convolutional neural network (CNN) on the area under the receiver operating characteristic curve (AUROC) at every training size, most widely with the fewest labeled cases. On two of the other three the CNN leads with the fewest labeled cases before the mechanistic model overtakes it. We report two experiments. On a 17-diagnosis PhysioNet task (5,108 recordings), adding an unconstrained language model to a gradient-boosted classifier lowers precision from 0.68 to 0.60 and micro-averaged F1 (micro-F1) from 0.72 to 0.69, whereas the same frozen model applying validated Skills reaches micro-F1 0.80 and 57% exact match, against 0.73 for a classifier retrained on the additional labeled data. In a scale-up to 32 diagnoses on 63,256 recordings drawn predominantly from MIMIC-IV-ECG, the classifier swarm reaches a mean AUROC of 0.950, and 0.929 against same-source negatives only. The agent runs fully offline on a single consumer laptop and supports interactive case review with the supervising clinician.

Zenodo (CERN European Organization for Nuclear Research)
University Children’s Hospital Basel (CH)
Quality Education
ECG Monitoring and Analysis
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.