A Controlled Single-Speaker Evaluation of Ambient Voice Technologies’ Speech-to-Text Function Using Narrated Surgical Case-Reports
Abstract Ambient voice technologies (AVT) utilise automatic speech recognition (ASR) and generative artificial intelligence for speech-to-text conversion, and are advocated as efficient, accurate tools for clinical documentation. We investigated ( n = 8) AVT systems using surgical case-reports as narratives in a controlled single-speaker evaluation prioritising safety outcomes: specifically 3 commercial transcription-systems (Dragon-Medical-One, Heidi-Health, Tortus); 4 ASR speech-to-text Application Programming Interfaces (Speechmatics-Enhanced, Amazon-Medical-Transcribe, Whisper, GPT4oTranscribe); and an experimental two-stage ASR–Large Language Model (LLM) pipeline incorporating GPT4oTranscribe with GPT-5-LLM generative error correction (GPT4oTranscribe-Corrected-5). Reference-transcripts ( n = 100; 32,897 words, range = 44–449, mean = 329/per-transcript) derived from surgical case-reports were recorded and input into AVT systems as audio-recordings for transcription-output generation. Primary outcome was proportion of transcription-outputs containing at least one clinically significant Class 3 error graded for potential harm. Secondary outcomes included transcription accuracy: Domain-Word-Error-Rate (DWER) against SNOMED-CT, lexical-accuracy (ROUGE score) and semantic similarity (BERT, BART scores). To investigate impact of the experimental pipeline on clinically significant errors, raw ASR and LLM-processed transcription-outputs were compared and errors classified as resolved, remaining or newly introduced. Across ( n = 800) transcript-outputs, 30–68% (GPT4oTranscribe-Corrected-5; Amazon-Medical-Transcribe and Dragon-Medical-One, respectively) contained at least one clinically significant Class 3 error. LLM-processing reduced the proportion of transcript-outputs affected (53% to 30%), amongst 89 clinically significant errors in ASR-output, 51 were resolved, 38 remained significant and 4 were newly introduced. Reference-transcript length significantly increased odds of Class 3 errors for 4 systems (GPT4oTranscribe, Tortus, Amazon-Medical-Transcribe, Dragon-Medical-One; odds ratio 2.65–3.34/additional 100 reference-words). Errors with potential to cause severe harm or death (NHS England Level 3–4) totalled ( n = 53) across systems, concentrated in the domains of medication-type/dose (39.6%), and investigations/laboratory results (30.2%). Performance varied across systems for transcription metrics ( P < 0.001). GPT4oTranscribe-Corrected-5 (DWER = 3.60%) and Heidi-Health (DWER = 5.67%) performed best for medical terminology; Amazon-Medical-Transcribe worst (DWER = 24.03%). For lexical and semantic similarity GPT4oTranscribe-Corrected-5 performed best, followed by Heidi-Health. The ASR–LLM pipeline reduced clinically significant errors but did introduce new ones in 3% of transcript-outputs. The safety and effectiveness of these systems in clinical practice requires further evaluation.
Authors
- Jadbinder Seehra (ORCID: https://orcid.org/0000-0002-3243-1580)
- Spyridon N. Papageorgiou (ORCID: https://orcid.org/0000-0003-1968-3326)
- Daniel Stonehouse‐Smith (ORCID: https://orcid.org/0000-0002-1096-5012)
- Melody Shirazi
- Martyn T. Cobourne (ORCID: https://orcid.org/0000-0003-2857-0315)
- Ruairí O’Kane (ORCID: https://orcid.org/0000-0002-7287-364X)
- Morgan Gregg
- Rachel Hutchinson
- Riya Patel
Institutions
- King's College London (GB)
- Guy's and St Thomas' NHS Foundation Trust (GB)
- University of Zurich (CH)
- Royal Alexandra Children's Hospital (GB)
- King's College Hospital NHS Foundation Trust (GB)
Publication Details
- Journal
- Journal of Medical Systems
- Published
- 2026-10-09
- DOI
- https://doi.org/10.1007/s10916-026-02471-5
- Primary Topic
- Electronic Health Records Systems
- Type
- article
- Field-Weighted Citation Impact
- 0.00