AI-Mediated Input Transformation in Medical English: Speech Recognition and Transcript Reliability
This study evaluates automatic speech recognition (ASR) performance on a controlled text-tospeech (TTS) rendering of a 2,692-token biochemistry passage used in Medical English instruction.Synthetic audio was transcribed by four online ASR platforms: Turboscribe, Riverside, 1transcribe, andSmartnote AI. Two freely accessible services (Turboscribe and Riverside) produced near-completetranscripts with 88.0–88.2% coverage and moderate Word Error Rates (WER) of 15.23% and 15.02%,respectively. Two paid services returned partial outputs covering 30.1% and 54.6% of the referencetext, yielding WER values of 72.14% and 47.77%, primarily driven by missing segments rather thandense substitution errors. Across systems, deletions were the dominant error type. In full transcripts,deletion rates were approximately 12.6% of reference tokens, with roughly 22% of deletions involvingstructural markers such as figure references and numeric labels. Meaning-altering semanticsubstitutions were rare (<1% of tokens) but pedagogically significant (e.g., imino → amino, cystine →cysteine, pH → phase). The findings demonstrate that coverage must be reported alongside WER inESP contexts, that full transcript export is essential for pedagogical reliability, and that lightweighthuman verification is required where terminological precision is critical.
Authors
- Evgeni Stanchev (ORCID: https://orcid.org/0009-0001-1108-1067)
Publication Details
- Published
- 2026-09-17
- DOI
- https://doi.org/10.35542/osf.io/f38mr_v1
- Primary Topic
- Artificial Intelligence in Healthcare and Education
- Type
- preprint