The Ammonix RCM Agent: Learning to Collect Healthcare Claims from Recorded Outcomes

After a medical service is delivered, getting it paid is a sequence of decisions: submit, correct, appeal, escalate to review, bill the balance, or write it off, and whether a choice was right is settled only later, by whether the payer pays. At one academic health system this work was measured at 14.5% of revenue for primary care visits and as much as 25.2% for emergency department visits. We ask whether a system that never lets a language model decide can outperform ones that do. The Ammonix RCM Agent, an instantiation of our companion architecture, decides from the recorded outcomes of past claims; frozen local language models write the forms and the rationale but never decide, and the agent hands a claim to a person when its evidence is thin. We evaluate it in a synthetic revenue-cycle world whose training corpus was generated by deliberately imperfect simulated billers and contains a planted trap where the common biller response is wrong, so the agent must learn beyond its teachers from logged outcomes alone. On three independent draws of 300 fresh claims, with every policy's paperwork written by the same model and judged by the same payer, the agent resolved 48.3% of claims on average against 45.8% for the strongest of five language-model baselines, the same open model given retrieval over the agent's own stored claims; it never finished behind on any draw against any baseline, won the claim-by-claim comparison against every baseline, and made no language-model call on its decision path; given GPT-6 as its writer, it matched or beat every configuration of GPT-6 deciding for itself at a quarter of the token cost. The same success signal trains the writer: rewarding drafts the payer accepted cut rejections of a 9B writer's appeal letters by nearly two thirds, below both a 27B and a frontier model. Because experience is stored and looked up rather than trained into weights, each settled claim informs the next decision with nothing retrained, and every decision can be traced to the stored claims that produced it.

Authors

Institutions

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-21
DOI
https://doi.org/10.5281/zenodo.22871211
Primary Topic
Artificial Intelligence in Healthcare and Education
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
preprint

The Ammonix RCM Agent: Learning to Collect Healthcare Claims from Recorded Outcomes

Francesca Stingele, Peter Ruppersberg, Duangjai Glauser, Lea Grieder et al.
Zenodo (CERN European Organization for Nuclear Research)
Artificial Intelligence in Healthcare and Education
preprint

The Ammonix RCM Agent: Learning to Collect Healthcare Claims from Recorded Outcomes

Francesca Stingele, Peter Ruppersberg, Duangjai Glauser, Lea Grieder, Adele Glauser, Matthew Todorov
preprint en

Abstract

After a medical service is delivered, getting it paid is a sequence of decisions: submit, correct, appeal, escalate to review, bill the balance, or write it off, and whether a choice was right is settled only later, by whether the payer pays. At one academic health system this work was measured at 14.5% of revenue for primary care visits and as much as 25.2% for emergency department visits. We ask whether a system that never lets a language model decide can outperform ones that do. The Ammonix RCM Agent, an instantiation of our companion architecture, decides from the recorded outcomes of past claims; frozen local language models write the forms and the rationale but never decide, and the agent hands a claim to a person when its evidence is thin. We evaluate it in a synthetic revenue-cycle world whose training corpus was generated by deliberately imperfect simulated billers and contains a planted trap where the common biller response is wrong, so the agent must learn beyond its teachers from logged outcomes alone. On three independent draws of 300 fresh claims, with every policy's paperwork written by the same model and judged by the same payer, the agent resolved 48.3% of claims on average against 45.8% for the strongest of five language-model baselines, the same open model given retrieval over the agent's own stored claims; it never finished behind on any draw against any baseline, won the claim-by-claim comparison against every baseline, and made no language-model call on its decision path; given GPT-6 as its writer, it matched or beat every configuration of GPT-6 deciding for itself at a quarter of the token cost. The same success signal trains the writer: rewarding drafts the payer accepted cut rejections of a 9B writer's appeal letters by nearly two thirds, below both a 27B and a frontier model. Because experience is stored and looked up rather than trained into weights, each settled claim informs the next decision with nothing retrained, and every decision can be traced to the stored claims that produced it.

Zenodo (CERN European Organization for Nuclear Research)
Amunix (United States) (US)
Peace, Justice and strong institutions
Artificial Intelligence in Healthcare and Education
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.