Semantic Segment Alignment and Granular ASR Performance Evaluation with Segment- and Word-Level WER

Word Error Rate (WER) is the de facto standard for evaluating automatic speech recognition (ASR) systems. Yet, conventional WER is a global average, which obscures critical performance variations across semantically important transcript segments; segment-level WER can reveal deeper insights for performance evaluation of ASR systems in many applications, and word-level WER can provide an additional perspective that can guide fine-tuning for ASR models. In this paper, we focus on segment- and word-level WER calculation. While it is possible to calculate these metrics using custom postprocessing scripts on top of existing tools such as SCTK/sclite, this approach has not been widely reported, likely due to limited application demand. Our work is motivated by ASR applications to air traffic control (ATC) communications, where safety-critical semantic segments—such as aircraft callsigns, action commands, and numerical values—demand prioritized performance monitoring. While evaluating the postprocessing-based approach, we also explored alternatives. Through systematic analysis of Levenshtein edit distance (LD) table patterns, we derive rules for aligning semantic segments in reference transcripts to corresponding hypothesis transcripts and computing localized—segment- and word-level—LD, used for evaluating local WER. Based on these rules, this paper introduces a comprehensive framework for segment alignment and segment- and word-level WER evaluation using dynamic programming algorithms that are markedly more efficient than postprocessing-based approaches using existing tools. We present three fast algorithms for segment alignment, as well as segment- and word-level LD calculations. To address accuracy issues caused by prefix drift (hallucination), we introduce refined implementations that mitigate this problem. For the word-level case, we introduce a generalized LD (GLD) with associated attribution rules that distribute shared blame to words potentially responsible for hallucinations. We validate the word-level GLD algorithm through simulation, where the ground-truth value for GLD attribution is obtained from perfect word-level transcript alignment established using simultaneously generated reference and hypothesis transcripts. All algorithms are implemented in an open-source Python tool called salwer, designed for voice-assisted avionics applications based on ATC voice instructions. The tool can be readily modified or extended to other domains requiring localized error evaluation. Experimental results demonstrate that segment- and word-level WER reveal critical insights hidden in global averages, particularly for high-risk words and segments. Analysis results based on our framework enable fine-grained ASR performance assessment and support targeted model improvements.

Authors

Institutions

Publication Details

Journal
Computers
Published
2026-10-09
DOI
https://doi.org/10.3390/computers15100691
Primary Topic
Speech Recognition and Synthesis
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

Semantic Segment Alignment and Granular ASR Performance Evaluation with Segment- and Word-Level WER

Jianhua Liu
Computers
Speech Recognition and Synthesis
article

Semantic Segment Alignment and Granular ASR Performance Evaluation with Segment- and Word-Level WER

Jianhua Liu
article en

Abstract

Word Error Rate (WER) is the de facto standard for evaluating automatic speech recognition (ASR) systems. Yet, conventional WER is a global average, which obscures critical performance variations across semantically important transcript segments; segment-level WER can reveal deeper insights for performance evaluation of ASR systems in many applications, and word-level WER can provide an additional perspective that can guide fine-tuning for ASR models. In this paper, we focus on segment- and word-level WER calculation. While it is possible to calculate these metrics using custom postprocessing scripts on top of existing tools such as SCTK/sclite, this approach has not been widely reported, likely due to limited application demand. Our work is motivated by ASR applications to air traffic control (ATC) communications, where safety-critical semantic segments—such as aircraft callsigns, action commands, and numerical values—demand prioritized performance monitoring. While evaluating the postprocessing-based approach, we also explored alternatives. Through systematic analysis of Levenshtein edit distance (LD) table patterns, we derive rules for aligning semantic segments in reference transcripts to corresponding hypothesis transcripts and computing localized—segment- and word-level—LD, used for evaluating local WER. Based on these rules, this paper introduces a comprehensive framework for segment alignment and segment- and word-level WER evaluation using dynamic programming algorithms that are markedly more efficient than postprocessing-based approaches using existing tools. We present three fast algorithms for segment alignment, as well as segment- and word-level LD calculations. To address accuracy issues caused by prefix drift (hallucination), we introduce refined implementations that mitigate this problem. For the word-level case, we introduce a generalized LD (GLD) with associated attribution rules that distribute shared blame to words potentially responsible for hallucinations. We validate the word-level GLD algorithm through simulation, where the ground-truth value for GLD attribution is obtained from perfect word-level transcript alignment established using simultaneously generated reference and hypothesis transcripts. All algorithms are implemented in an open-source Python tool called salwer, designed for voice-assisted avionics applications based on ATC voice instructions. The tool can be readily modified or extended to other domains requiring localized error evaluation. Experimental results demonstrate that segment- and word-level WER reveal critical insights hidden in global averages, particularly for high-risk words and segments. Analysis results based on our framework enable fine-grained ASR performance assessment and support targeted model improvements.

ComputersVol. 15(10)
Embry–Riddle Aeronautical University (US)
Openalex Percentile: Top 12%
Speech Recognition and Synthesis
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.