Similarity-Aware Viseme Caching for Efficient Lip-Synchronization of English Speech

Speech-driven lip-synchronization systems are increasingly used in virtual avatars, digital assistants, and interactive media. Although recent models can generate realistic talking-head videos, visual synthesis remains computationally expensive, particularly when large numbers of short utterances must be produced in real time. In practical deployments, many generated segments contain perceptually similar mouth movements even when the input text differs, yet current pipelines regenerate lip motion for every request because no mechanism exists to safely reuse previously generated animations. This study proposes a similarity-aware viseme caching framework that reduces redundant visual generation for English speech. Instead of matching input text or audio, the method operates on viseme sequences representing visually distinguishable mouth configurations derived from phoneme-to-viseme mapping. A composite similarity metric combines weighted edit distance, dynamic time warping, longest common subsequence, and local positional and transition features using a perceptually motivated viseme cost matrix. An offline optimization procedure aligns similarity scores with perceptual acceptability and enables threshold-based reuse decisions. Given a new utterance, its viseme sequence is compared against previously generated segments. If similarity exceeds a predefined threshold, the stored video is reused; otherwise, normal GPU-based lip-sync generation is performed, and the result is cached. The similarity search over a 1002-video database requires approximately 50–100 ms on a 16-core CPU, enabling real-time deployment. Deployment on a live system with 7000 cached videos handling 6250 user requests achieved a cache hit rate and GPU utilization savings of $$\\approx$$ 83% at similarity threshold $$\\theta$$ = 0.60, reducing the average generation time by $$\\approx$$ 7.5 seconds per request, for this deployment’s short-form, repetitive-utterance workload. Reuse decisions were calibrated against human perceptual ratings of viseme-sequence similarity; this study did not include a direct paired comparison (e.g., a 2AFC study) between cached and freshly generated videos, and the reported hit rate reflects this specific workload rather than open-ended conversational use. The results demonstrate that perceptually calibrated sequence similarity can transform lip-sync generation from a purely generative process into a reusable systems pipeline, substantially improving computational efficiency for short, repetitive-utterance workloads. Direct perceptual comparison of cached versus freshly generated video, and evaluation on more open-ended conversational workloads, remain important directions for future validation.

Authors

Institutions

Publication Details

Journal
Human-Centric Intelligent Systems
Published
2026-08-25
DOI
https://doi.org/10.1007/s44230-026-00170-5
Primary Topic
Speech and Audio Processing
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Similarity-Aware Viseme Caching for Efficient Lip-Synchronization of English Speech

Deniz Kenan Kılıç, Bahadir İrfan Katıpoğlu, Mehmet Salih Zeman
Human-Centric Intelligent Systems
Speech and Audio Processing
article

Similarity-Aware Viseme Caching for Efficient Lip-Synchronization of English Speech

Deniz Kenan Kılıç, Bahadir İrfan Katıpoğlu, Mehmet Salih Zeman
article en

Abstract

Speech-driven lip-synchronization systems are increasingly used in virtual avatars, digital assistants, and interactive media. Although recent models can generate realistic talking-head videos, visual synthesis remains computationally expensive, particularly when large numbers of short utterances must be produced in real time. In practical deployments, many generated segments contain perceptually similar mouth movements even when the input text differs, yet current pipelines regenerate lip motion for every request because no mechanism exists to safely reuse previously generated animations. This study proposes a similarity-aware viseme caching framework that reduces redundant visual generation for English speech. Instead of matching input text or audio, the method operates on viseme sequences representing visually distinguishable mouth configurations derived from phoneme-to-viseme mapping. A composite similarity metric combines weighted edit distance, dynamic time warping, longest common subsequence, and local positional and transition features using a perceptually motivated viseme cost matrix. An offline optimization procedure aligns similarity scores with perceptual acceptability and enables threshold-based reuse decisions. Given a new utterance, its viseme sequence is compared against previously generated segments. If similarity exceeds a predefined threshold, the stored video is reused; otherwise, normal GPU-based lip-sync generation is performed, and the result is cached. The similarity search over a 1002-video database requires approximately 50–100 ms on a 16-core CPU, enabling real-time deployment. Deployment on a live system with 7000 cached videos handling 6250 user requests achieved a cache hit rate and GPU utilization savings of $$\approx$$ 83% at similarity threshold $$\theta$$ = 0.60, reducing the average generation time by $$\approx$$ 7.5 seconds per request, for this deployment’s short-form, repetitive-utterance workload. Reuse decisions were calibrated against human perceptual ratings of viseme-sequence similarity; this study did not include a direct paired comparison (e.g., a 2AFC study) between cached and freshly generated videos, and the reported hit rate reflects this specific workload rather than open-ended conversational use. The results demonstrate that perceptually calibrated sequence similarity can transform lip-sync generation from a purely generative process into a reusable systems pipeline, substantially improving computational efficiency for short, repetitive-utterance workloads. Direct perceptual comparison of cached versus freshly generated video, and evaluation on more open-ended conversational workloads, remain important directions for future validation.

Human-Centric Intelligent Systems
Technopolis (Finland) (FI)
Openalex Percentile: Top 9%
Speech and Audio Processing
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.