Asking Instead of Listening: A Semantic Voice Activity Detector for Turn Decisions
While acoustic Voice Activity Detection (VAD) determines the presence of speech, conversational voice pipelines require semantic decision-making to handle turn-taking dynamics. This paper introduces jev-vad, a lightweight, model-agnostic semantic VAD layer that replaces heuristic barge-in and end-of-turn predictions with typed questions evaluated by a small System-1 decision model. Instantiated via an int4-quantized ONNX export of Laya (a ~206 MB, CPU-only mmBERT-base encoder), jev-vad processes serialized dialogue state (the agent's prior utterance plus the partial transcript) in a single non-autoregressive forward pass to yield per-option choice probabilities. We demonstrate that barge-in and turn-end decisions possess opposing failure-cost asymmetries, requiring two separately tuned questions: barge-in detection requires recall-biased tuning (achieving an F1 score of 0.76 at a 0.30 threshold with a filler guard), whereas end-of-turn prediction demands precision-biased tuning combined with acoustic pause detection. Evaluated on a 100-phrase Spanish benchmark, we uncover a language-mixing phenomenon where English question framing over Spanish dialogue state yields up to a +0.40 accuracy improvement for barge-in detection, while turn-end detection requires native Spanish framing. The winning configuration runs in 101–187 ms p50 on an 8-vCPU cloud instance, providing a fully self-hosted solution alongside a reproducible ablation methodology for cross-lingual adaptation.
Authors
- Daniel García (ORCID: https://orcid.org/0000-0002-1456-4785)
Institutions
- Universidad de Oviedo (ES)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-10-08
- DOI
- https://doi.org/10.5281/zenodo.23233330
- Primary Topic
- Speech and dialogue systems
- Type
- article
- Field-Weighted Citation Impact
- 0.00