Similarity-Aware Viseme Caching for Efficient Lip-Synchronization of English Speech
Speech-driven lip-synchronization systems are increasingly used in virtual avatars, digital assistants, and interactive media. Although recent models can generate realistic talking-head videos, visual synthesis remains computationally expensive, particularly when large numbers of short utterances must be produced in real time. In practical deployments, many generated segments contain perceptually similar mouth movements even when the input text differs, yet current pipelines regenerate lip motion for every request because no mechanism exists to safely reuse previously generated animations. This study proposes a similarity-aware viseme caching framework that reduces redundant visual generation for English speech. Instead of matching input text or audio, the method operates on viseme sequences representing visually distinguishable mouth configurations derived from phoneme-to-viseme mapping. A composite similarity metric combines weighted edit distance, dynamic time warping, longest common subsequence, and local positional and transition features using a perceptually motivated viseme cost matrix. An offline optimization procedure aligns similarity scores with perceptual acceptability and enables threshold-based reuse decisions. Given a new utterance, its viseme sequence is compared against previously generated segments. If similarity exceeds a predefined threshold, the stored video is reused; otherwise, normal GPU-based lip-sync generation is performed, and the result is cached. The similarity search over a 1002-video database requires approximately 50–100 ms on a 16-core CPU, enabling real-time deployment. Deployment on a live system with 7000 cached videos handling 6250 user requests achieved a cache hit rate and GPU utilization savings of $$\\approx$$ 83% at similarity threshold $$\\theta$$ = 0.60, reducing the average generation time by $$\\approx$$ 7.5 seconds per request, for this deployment’s short-form, repetitive-utterance workload. Reuse decisions were calibrated against human perceptual ratings of viseme-sequence similarity; this study did not include a direct paired comparison (e.g., a 2AFC study) between cached and freshly generated videos, and the reported hit rate reflects this specific workload rather than open-ended conversational use. The results demonstrate that perceptually calibrated sequence similarity can transform lip-sync generation from a purely generative process into a reusable systems pipeline, substantially improving computational efficiency for short, repetitive-utterance workloads. Direct perceptual comparison of cached versus freshly generated video, and evaluation on more open-ended conversational workloads, remain important directions for future validation.
Authors
- Deniz Kenan Kılıç (ORCID: https://orcid.org/0000-0002-6996-3425)
- Bahadir İrfan Katıpoğlu
- Mehmet Salih Zeman (ORCID: https://orcid.org/0009-0003-3911-6828)
Institutions
- Technopolis (Finland) (FI)
Publication Details
- Journal
- Human-Centric Intelligent Systems
- Published
- 2026-08-25
- DOI
- https://doi.org/10.1007/s44230-026-00170-5
- Primary Topic
- Speech and Audio Processing
- Type
- article
- Field-Weighted Citation Impact
- 0.00