RooM: Real-Time Event-Driven Streaming Architecture for Vietnamese Voice Call Meeting AI
Modern enterprise knowledge management relies increasingly on automated conversational intelligence to capture decisions and summarize collaborative conferences. However, deploying interactive voice-driven assistants into live virtual meetings (e.g., Google Meet, MS Teams, Zoom) presents severe technical bottlenecks: high acoustic echo, background noise, tonal linguistic ambiguity in Vietnamese, context window degradation in 120-minute sessions, and strict hardware ceilings on consumer edge devices. This paper introduces RooM (Voice Call Meeting AI), a novel event-driven streaming architecture engineered for real-time Vietnamese meeting intelligence on standard consumer hardware equipped with an Intel Core i7-12650H CPU and an NVIDIA GeForce RTX 3050 Laptop GPU (4GB VRAM). RooM establishes a non-blocking asynchronous pipeline structured across seven decoupled modules: (1) 3-tier acoustic signal enhancement via WebRTC APM, DeepFilterNet3, and Silero VAD v5 (256ms hangover); (2) low-power wake-word detection using OpenWakeWord with temporal moving average smoothing; (3) streaming Vietnamese ASR using Faster-Whisper CTranslate2 INT8 coupled with Local Agreement Prefix Alignment; (4) joint intent classification and slot filling using a fine-tuned PhoBERT-Joint core with ISO-8601 canonicalization; (5) an OS-inspired 3-tier virtual memory hierarchy (MemGPT architecture) with recency decay scheduling resolving context window limits under a 3000-token budget; (6) a Dense-Sparse Hybrid RAG engine (BGE-M3+ Okapi BM25 with Reciprocal Rank Fusion) coupled with a quantized Qwen2.5-3B-Instruct core under Context-Free Grammar logit masking; and (7) a dual-path multimodal streamer delivering WebSocket text transcripts and neural synthetic speech via Kokoro-82M. Empirical evaluations demonstrate that RooM achieves an acoustic SNR gain of +14.6 dB, an ASR Word Error Rate (WER) of 8.6%, an Intent F1 of 93.8%, a Slot F1 of 91.2%, and an end-to-end voice response latency under 850 ms (audio playback starts at 640 ms) while capping peak GPU VRAM at 3.15 GB, entirely preventing CUDA Out-Of-Memory exceptions without reliance on external cloud APIs.
Authors
- Ngoc Minh Vu
Institutions
- Industrial University of Ho Chi Minh City (VN)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-15
- DOI
- https://doi.org/10.5281/zenodo.22761611
- Primary Topic
- Speech Recognition and Synthesis
- Type
- preprint