A Controlled Single-Speaker Evaluation of Ambient Voice Technologies’ Speech-to-Text Function Using Narrated Surgical Case-Reports

Abstract Ambient voice technologies (AVT) utilise automatic speech recognition (ASR) and generative artificial intelligence for speech-to-text conversion, and are advocated as efficient, accurate tools for clinical documentation. We investigated ( n = 8) AVT systems using surgical case-reports as narratives in a controlled single-speaker evaluation prioritising safety outcomes: specifically 3 commercial transcription-systems (Dragon-Medical-One, Heidi-Health, Tortus); 4 ASR speech-to-text Application Programming Interfaces (Speechmatics-Enhanced, Amazon-Medical-Transcribe, Whisper, GPT4oTranscribe); and an experimental two-stage ASR–Large Language Model (LLM) pipeline incorporating GPT4oTranscribe with GPT-5-LLM generative error correction (GPT4oTranscribe-Corrected-5). Reference-transcripts ( n = 100; 32,897 words, range = 44–449, mean = 329/per-transcript) derived from surgical case-reports were recorded and input into AVT systems as audio-recordings for transcription-output generation. Primary outcome was proportion of transcription-outputs containing at least one clinically significant Class 3 error graded for potential harm. Secondary outcomes included transcription accuracy: Domain-Word-Error-Rate (DWER) against SNOMED-CT, lexical-accuracy (ROUGE score) and semantic similarity (BERT, BART scores). To investigate impact of the experimental pipeline on clinically significant errors, raw ASR and LLM-processed transcription-outputs were compared and errors classified as resolved, remaining or newly introduced. Across ( n = 800) transcript-outputs, 30–68% (GPT4oTranscribe-Corrected-5; Amazon-Medical-Transcribe and Dragon-Medical-One, respectively) contained at least one clinically significant Class 3 error. LLM-processing reduced the proportion of transcript-outputs affected (53% to 30%), amongst 89 clinically significant errors in ASR-output, 51 were resolved, 38 remained significant and 4 were newly introduced. Reference-transcript length significantly increased odds of Class 3 errors for 4 systems (GPT4oTranscribe, Tortus, Amazon-Medical-Transcribe, Dragon-Medical-One; odds ratio 2.65–3.34/additional 100 reference-words). Errors with potential to cause severe harm or death (NHS England Level 3–4) totalled ( n = 53) across systems, concentrated in the domains of medication-type/dose (39.6%), and investigations/laboratory results (30.2%). Performance varied across systems for transcription metrics ( P < 0.001). GPT4oTranscribe-Corrected-5 (DWER = 3.60%) and Heidi-Health (DWER = 5.67%) performed best for medical terminology; Amazon-Medical-Transcribe worst (DWER = 24.03%). For lexical and semantic similarity GPT4oTranscribe-Corrected-5 performed best, followed by Heidi-Health. The ASR–LLM pipeline reduced clinically significant errors but did introduce new ones in 3% of transcript-outputs. The safety and effectiveness of these systems in clinical practice requires further evaluation.

Authors

Institutions

Publication Details

Journal
Journal of Medical Systems
Published
2026-10-09
DOI
https://doi.org/10.1007/s10916-026-02471-5
Primary Topic
Electronic Health Records Systems
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

A Controlled Single-Speaker Evaluation of Ambient Voice Technologies’ Speech-to-Text Function Using Narrated Surgical Case-Reports

Jadbinder Seehra, Spyridon N. Papageorgiou, Daniel Stonehouse‐Smith, Melody Shirazi et al.
Journal of Medical Systems
Electronic Health Records Systems
article

A Controlled Single-Speaker Evaluation of Ambient Voice Technologies’ Speech-to-Text Function Using Narrated Surgical Case-Reports

Jadbinder Seehra, Spyridon N. Papageorgiou, Daniel Stonehouse‐Smith, Melody Shirazi, Martyn T. Cobourne, Ruairí O’Kane, Morgan Gregg, Rachel Hutchinson, Riya Patel
article en

Abstract

Abstract Ambient voice technologies (AVT) utilise automatic speech recognition (ASR) and generative artificial intelligence for speech-to-text conversion, and are advocated as efficient, accurate tools for clinical documentation. We investigated ( n = 8) AVT systems using surgical case-reports as narratives in a controlled single-speaker evaluation prioritising safety outcomes: specifically 3 commercial transcription-systems (Dragon-Medical-One, Heidi-Health, Tortus); 4 ASR speech-to-text Application Programming Interfaces (Speechmatics-Enhanced, Amazon-Medical-Transcribe, Whisper, GPT4oTranscribe); and an experimental two-stage ASR–Large Language Model (LLM) pipeline incorporating GPT4oTranscribe with GPT-5-LLM generative error correction (GPT4oTranscribe-Corrected-5). Reference-transcripts ( n = 100; 32,897 words, range = 44–449, mean = 329/per-transcript) derived from surgical case-reports were recorded and input into AVT systems as audio-recordings for transcription-output generation. Primary outcome was proportion of transcription-outputs containing at least one clinically significant Class 3 error graded for potential harm. Secondary outcomes included transcription accuracy: Domain-Word-Error-Rate (DWER) against SNOMED-CT, lexical-accuracy (ROUGE score) and semantic similarity (BERT, BART scores). To investigate impact of the experimental pipeline on clinically significant errors, raw ASR and LLM-processed transcription-outputs were compared and errors classified as resolved, remaining or newly introduced. Across ( n = 800) transcript-outputs, 30–68% (GPT4oTranscribe-Corrected-5; Amazon-Medical-Transcribe and Dragon-Medical-One, respectively) contained at least one clinically significant Class 3 error. LLM-processing reduced the proportion of transcript-outputs affected (53% to 30%), amongst 89 clinically significant errors in ASR-output, 51 were resolved, 38 remained significant and 4 were newly introduced. Reference-transcript length significantly increased odds of Class 3 errors for 4 systems (GPT4oTranscribe, Tortus, Amazon-Medical-Transcribe, Dragon-Medical-One; odds ratio 2.65–3.34/additional 100 reference-words). Errors with potential to cause severe harm or death (NHS England Level 3–4) totalled ( n = 53) across systems, concentrated in the domains of medication-type/dose (39.6%), and investigations/laboratory results (30.2%). Performance varied across systems for transcription metrics ( P < 0.001). GPT4oTranscribe-Corrected-5 (DWER = 3.60%) and Heidi-Health (DWER = 5.67%) performed best for medical terminology; Amazon-Medical-Transcribe worst (DWER = 24.03%). For lexical and semantic similarity GPT4oTranscribe-Corrected-5 performed best, followed by Heidi-Health. The ASR–LLM pipeline reduced clinically significant errors but did introduce new ones in 3% of transcript-outputs. The safety and effectiveness of these systems in clinical practice requires further evaluation.

Journal of Medical SystemsVol. 50(1)
King's College London (GB), Guy's and St Thomas' NHS Foundation Trust (GB), University of Zurich (CH), Royal Alexandra Children's Hospital (GB), King's College Hospital NHS Foundation Trust (GB)
Openalex Percentile: Top 8%
Electronic Health Records Systems
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.