Semantic Segment Alignment and Granular ASR Performance Evaluation with Segment- and Word-Level WER
Word Error Rate (WER) is the de facto standard for evaluating automatic speech recognition (ASR) systems. Yet, conventional WER is a global average, which obscures critical performance variations across semantically important transcript segments; segment-level WER can reveal deeper insights for performance evaluation of ASR systems in many applications, and word-level WER can provide an additional perspective that can guide fine-tuning for ASR models. In this paper, we focus on segment- and word-level WER calculation. While it is possible to calculate these metrics using custom postprocessing scripts on top of existing tools such as SCTK/sclite, this approach has not been widely reported, likely due to limited application demand. Our work is motivated by ASR applications to air traffic control (ATC) communications, where safety-critical semantic segments—such as aircraft callsigns, action commands, and numerical values—demand prioritized performance monitoring. While evaluating the postprocessing-based approach, we also explored alternatives. Through systematic analysis of Levenshtein edit distance (LD) table patterns, we derive rules for aligning semantic segments in reference transcripts to corresponding hypothesis transcripts and computing localized—segment- and word-level—LD, used for evaluating local WER. Based on these rules, this paper introduces a comprehensive framework for segment alignment and segment- and word-level WER evaluation using dynamic programming algorithms that are markedly more efficient than postprocessing-based approaches using existing tools. We present three fast algorithms for segment alignment, as well as segment- and word-level LD calculations. To address accuracy issues caused by prefix drift (hallucination), we introduce refined implementations that mitigate this problem. For the word-level case, we introduce a generalized LD (GLD) with associated attribution rules that distribute shared blame to words potentially responsible for hallucinations. We validate the word-level GLD algorithm through simulation, where the ground-truth value for GLD attribution is obtained from perfect word-level transcript alignment established using simultaneously generated reference and hypothesis transcripts. All algorithms are implemented in an open-source Python tool called salwer, designed for voice-assisted avionics applications based on ATC voice instructions. The tool can be readily modified or extended to other domains requiring localized error evaluation. Experimental results demonstrate that segment- and word-level WER reveal critical insights hidden in global averages, particularly for high-risk words and segments. Analysis results based on our framework enable fine-grained ASR performance assessment and support targeted model improvements.
Authors
- Jianhua Liu (ORCID: https://orcid.org/0009-0004-9460-1495)
Institutions
- Embry–Riddle Aeronautical University (US)
Publication Details
- Journal
- Computers
- Published
- 2026-10-09
- DOI
- https://doi.org/10.3390/computers15100691
- Primary Topic
- Speech Recognition and Synthesis
- Type
- article
- Field-Weighted Citation Impact
- 0.00