AI-Mediated Input Transformation in Medical English: Speech Recognition and Transcript Reliability

This study evaluates automatic speech recognition (ASR) performance on a controlled text-tospeech (TTS) rendering of a 2,692-token biochemistry passage used in Medical English instruction.Synthetic audio was transcribed by four online ASR platforms: Turboscribe, Riverside, 1transcribe, andSmartnote AI. Two freely accessible services (Turboscribe and Riverside) produced near-completetranscripts with 88.0–88.2% coverage and moderate Word Error Rates (WER) of 15.23% and 15.02%,respectively. Two paid services returned partial outputs covering 30.1% and 54.6% of the referencetext, yielding WER values of 72.14% and 47.77%, primarily driven by missing segments rather thandense substitution errors. Across systems, deletions were the dominant error type. In full transcripts,deletion rates were approximately 12.6% of reference tokens, with roughly 22% of deletions involvingstructural markers such as figure references and numeric labels. Meaning-altering semanticsubstitutions were rare (<1% of tokens) but pedagogically significant (e.g., imino → amino, cystine →cysteine, pH → phase). The findings demonstrate that coverage must be reported alongside WER inESP contexts, that full transcript export is essential for pedagogical reliability, and that lightweighthuman verification is required where terminological precision is critical.

Authors

Publication Details

Published
2026-09-17
DOI
https://doi.org/10.35542/osf.io/f38mr_v1
Primary Topic
Artificial Intelligence in Healthcare and Education
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
preprint

AI-Mediated Input Transformation in Medical English: Speech Recognition and Transcript Reliability

Evgeni Stanchev
Artificial Intelligence in Healthcare and Education
preprint

AI-Mediated Input Transformation in Medical English: Speech Recognition and Transcript Reliability

Evgeni Stanchev
preprint en

Abstract

This study evaluates automatic speech recognition (ASR) performance on a controlled text-tospeech (TTS) rendering of a 2,692-token biochemistry passage used in Medical English instruction.Synthetic audio was transcribed by four online ASR platforms: Turboscribe, Riverside, 1transcribe, andSmartnote AI. Two freely accessible services (Turboscribe and Riverside) produced near-completetranscripts with 88.0–88.2% coverage and moderate Word Error Rates (WER) of 15.23% and 15.02%,respectively. Two paid services returned partial outputs covering 30.1% and 54.6% of the referencetext, yielding WER values of 72.14% and 47.77%, primarily driven by missing segments rather thandense substitution errors. Across systems, deletions were the dominant error type. In full transcripts,deletion rates were approximately 12.6% of reference tokens, with roughly 22% of deletions involvingstructural markers such as figure references and numeric labels. Meaning-altering semanticsubstitutions were rare (<1% of tokens) but pedagogically significant (e.g., imino → amino, cystine →cysteine, pH → phase). The findings demonstrate that coverage must be reported alongside WER inESP contexts, that full transcript export is essential for pedagogical reliability, and that lightweighthuman verification is required where terminological precision is critical.

Artificial Intelligence in Healthcare and Education
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

AI-Mediated Input Transformation in Medical English: Speech Recognition and Transcript Reliability — Evgeni Stanchev · (2026) | TGRS Research Map | TGRS