Audience Reception of Simultaneous Speech-to-Text Translation Subtitles: A Mixed-Methods Study
Simultaneous speech-to-text translation (SimulST) is increasingly used in multilingual settings, yet evaluation remains largely system-centered, with limited attention to user experience. This study examined audience reception of two commercial SimulST systems with different updating behaviors across English- and Japanese-source videos. Forty Mandarin-speaking university students viewed two TED/TEDx-style videos under one assigned system condition. Reception was assessed through questionnaires, eye tracking, and supplementary checks of subtitle quality, latency, and update behavior. System condition did not significantly affect comprehension, perceived translation quality, satisfaction, or self-reported cognitive load, but was associated with perceived synchronization and visual attention to subtitles. Source-language condition showed broader effects on comprehension, perceived quality, cognitive load, and subtitle reliance, while interaction effects emerged only for synchronization and distraction. Overall satisfaction remained modest, highlighting concerns about accuracy, latency, instability, and coherence, and the need for more audience-centered, interface-oriented SimulST evaluation.
Authors
- Yu Zhou (ORCID: https://orcid.org/0009-0001-1522-4846)
- Fanglu Xie
- Yang Zhou
Institutions
- Lingnan University (HK)
Publication Details
- Journal
- International Journal of Human-Computer Interaction
- Published
- 2026-09-25
- DOI
- https://doi.org/10.1080/10447318.2026.2734668
- Primary Topic
- Subtitles and Audiovisual Media
- Type
- article
- Field-Weighted Citation Impact
- 0.00