When the scribe does the reasoning: ambient artificial intelligence, inference impersonation, and the development of trainees’ clinical judgment

Ambient artificial intelligence (AI) scribes are entering clinical practice on aligned incentives, with physicians gaining relief from documentation burden and health systems gaining revenue through more complete coding. Yet adoption has outpaced evaluation. This paper identifies a specific phenomenon, which it calls inference impersonation. In the Assessment and Plan, ambient AI scribes do not transcribe clinical reasoning but rather generate it. The resulting text mimics physician judgment without the iterative hypothesis testing, contextual weighing, and management reasoning that define clinical thought. Generated and transcribed content are effectively indistinguishable in the final note, so neither the signing physician nor the reviewing trainee can identify where transcription ends and generation begins. The risk falls hardest on trainees. An experienced attending may recognize a narrow differential or a generic plan, but a trainee cannot evaluate what they cannot yet produce independently, and automation bias is powerful. In one randomized trial, physicians shown deliberately erroneous AI suggestions after a 20-hour artificial intelligence literacy program scored 73% on diagnostic reasoning versus 85% among those receiving error-free suggestions. When trainees learn to reason by editing AI drafts rather than producing their own, they risk never developing the judgment that training exists to build. The authors propose a response that is both technical and educational. Vendors should offer centrally managed, learner-specific configurations and section-level transparency distinguishing transcribed from generated content, and training programs should make these an explicit condition of deployment. Because no major commercial scribe today constrains inference for trainees, programs can protect an existing norm now by requiring trainees to present their reasoning before the note is generated, gating scribe access to demonstrated competency, and assessing reasoning through the oral case presentation as well as the written note, mapped to an entrustable activity and its milestones.

Authors

Institutions

Publication Details

Journal
Academic Medicine
Published
2026-09-15
DOI
https://doi.org/10.1093/acamed/wvag287
Primary Topic
Artificial Intelligence in Healthcare and Education
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

When the scribe does the reasoning: ambient artificial intelligence, inference impersonation, and the development of trainees’ clinical judgment

Robert L. Cloutier, John Lee, Jordan Wackett, Jane Abernethy et al.
Academic Medicine
Artificial Intelligence in Healthcare and Education
article

When the scribe does the reasoning: ambient artificial intelligence, inference impersonation, and the development of trainees’ clinical judgment

Robert L. Cloutier, John Lee, Jordan Wackett, Jane Abernethy, Steven McGaughey, R Logan Jones
article en

Abstract

Ambient artificial intelligence (AI) scribes are entering clinical practice on aligned incentives, with physicians gaining relief from documentation burden and health systems gaining revenue through more complete coding. Yet adoption has outpaced evaluation. This paper identifies a specific phenomenon, which it calls inference impersonation. In the Assessment and Plan, ambient AI scribes do not transcribe clinical reasoning but rather generate it. The resulting text mimics physician judgment without the iterative hypothesis testing, contextual weighing, and management reasoning that define clinical thought. Generated and transcribed content are effectively indistinguishable in the final note, so neither the signing physician nor the reviewing trainee can identify where transcription ends and generation begins. The risk falls hardest on trainees. An experienced attending may recognize a narrow differential or a generic plan, but a trainee cannot evaluate what they cannot yet produce independently, and automation bias is powerful. In one randomized trial, physicians shown deliberately erroneous AI suggestions after a 20-hour artificial intelligence literacy program scored 73% on diagnostic reasoning versus 85% among those receiving error-free suggestions. When trainees learn to reason by editing AI drafts rather than producing their own, they risk never developing the judgment that training exists to build. The authors propose a response that is both technical and educational. Vendors should offer centrally managed, learner-specific configurations and section-level transparency distinguishing transcribed from generated content, and training programs should make these an explicit condition of deployment. Because no major commercial scribe today constrains inference for trainees, programs can protect an existing norm now by requiring trainees to present their reasoning before the note is generated, gating scribe access to demonstrated competency, and assessing reasoning through the oral case presentation as well as the written note, mapped to an entrustable activity and its milestones.

Academic Medicine
Johns Hopkins University (US), Oregon Health & Science University (US), Johns Hopkins Medicine (US), Oregon Health and Science University Hospital (US)
Quality Education
Openalex Percentile: Top 15%
Artificial Intelligence in Healthcare and Education
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.