RooM: Real-Time Event-Driven Streaming Architecture for Vietnamese Voice Call Meeting AI

Modern enterprise knowledge management relies increasingly on automated conversational intelligence to capture decisions and summarize collaborative conferences. However, deploying interactive voice-driven assistants into live virtual meetings (e.g., Google Meet, MS Teams, Zoom) presents severe technical bottlenecks: high acoustic echo, background noise, tonal linguistic ambiguity in Vietnamese, context window degradation in 120-minute sessions, and strict hardware ceilings on consumer edge devices. This paper introduces RooM (Voice Call Meeting AI), a novel event-driven streaming architecture engineered for real-time Vietnamese meeting intelligence on standard consumer hardware equipped with an Intel Core i7-12650H CPU and an NVIDIA GeForce RTX 3050 Laptop GPU (4GB VRAM). RooM establishes a non-blocking asynchronous pipeline structured across seven decoupled modules: (1) 3-tier acoustic signal enhancement via WebRTC APM, DeepFilterNet3, and Silero VAD v5 (256ms hangover); (2) low-power wake-word detection using OpenWakeWord with temporal moving average smoothing; (3) streaming Vietnamese ASR using Faster-Whisper CTranslate2 INT8 coupled with Local Agreement Prefix Alignment; (4) joint intent classification and slot filling using a fine-tuned PhoBERT-Joint core with ISO-8601 canonicalization; (5) an OS-inspired 3-tier virtual memory hierarchy (MemGPT architecture) with recency decay scheduling resolving context window limits under a 3000-token budget; (6) a Dense-Sparse Hybrid RAG engine (BGE-M3+ Okapi BM25 with Reciprocal Rank Fusion) coupled with a quantized Qwen2.5-3B-Instruct core under Context-Free Grammar logit masking; and (7) a dual-path multimodal streamer delivering WebSocket text transcripts and neural synthetic speech via Kokoro-82M. Empirical evaluations demonstrate that RooM achieves an acoustic SNR gain of +14.6 dB, an ASR Word Error Rate (WER) of 8.6%, an Intent F1 of 93.8%, a Slot F1 of 91.2%, and an end-to-end voice response latency under 850 ms (audio playback starts at 640 ms) while capping peak GPU VRAM at 3.15 GB, entirely preventing CUDA Out-Of-Memory exceptions without reliance on external cloud APIs.

Authors

Institutions

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-15
DOI
https://doi.org/10.5281/zenodo.22761612
Primary Topic
Speech Recognition and Synthesis
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
preprint

RooM: Real-Time Event-Driven Streaming Architecture for Vietnamese Voice Call Meeting AI

Ngoc Minh Vu
Zenodo (CERN European Organization for Nuclear Research)
Speech Recognition and Synthesis
preprint

RooM: Real-Time Event-Driven Streaming Architecture for Vietnamese Voice Call Meeting AI

Ngoc Minh Vu
preprint en

Abstract

Modern enterprise knowledge management relies increasingly on automated conversational intelligence to capture decisions and summarize collaborative conferences. However, deploying interactive voice-driven assistants into live virtual meetings (e.g., Google Meet, MS Teams, Zoom) presents severe technical bottlenecks: high acoustic echo, background noise, tonal linguistic ambiguity in Vietnamese, context window degradation in 120-minute sessions, and strict hardware ceilings on consumer edge devices. This paper introduces RooM (Voice Call Meeting AI), a novel event-driven streaming architecture engineered for real-time Vietnamese meeting intelligence on standard consumer hardware equipped with an Intel Core i7-12650H CPU and an NVIDIA GeForce RTX 3050 Laptop GPU (4GB VRAM). RooM establishes a non-blocking asynchronous pipeline structured across seven decoupled modules: (1) 3-tier acoustic signal enhancement via WebRTC APM, DeepFilterNet3, and Silero VAD v5 (256ms hangover); (2) low-power wake-word detection using OpenWakeWord with temporal moving average smoothing; (3) streaming Vietnamese ASR using Faster-Whisper CTranslate2 INT8 coupled with Local Agreement Prefix Alignment; (4) joint intent classification and slot filling using a fine-tuned PhoBERT-Joint core with ISO-8601 canonicalization; (5) an OS-inspired 3-tier virtual memory hierarchy (MemGPT architecture) with recency decay scheduling resolving context window limits under a 3000-token budget; (6) a Dense-Sparse Hybrid RAG engine (BGE-M3+ Okapi BM25 with Reciprocal Rank Fusion) coupled with a quantized Qwen2.5-3B-Instruct core under Context-Free Grammar logit masking; and (7) a dual-path multimodal streamer delivering WebSocket text transcripts and neural synthetic speech via Kokoro-82M. Empirical evaluations demonstrate that RooM achieves an acoustic SNR gain of +14.6 dB, an ASR Word Error Rate (WER) of 8.6%, an Intent F1 of 93.8%, a Slot F1 of 91.2%, and an end-to-end voice response latency under 850 ms (audio playback starts at 640 ms) while capping peak GPU VRAM at 3.15 GB, entirely preventing CUDA Out-Of-Memory exceptions without reliance on external cloud APIs.

Zenodo (CERN European Organization for Nuclear Research)
Industrial University of Ho Chi Minh City (VN)
Quality Education
Speech Recognition and Synthesis
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

RooM: Real-Time Event-Driven Streaming Architecture for Vietnamese Voice Call Meeting AI — Ngoc Minh Vu · Zenodo (CERN European Organization for Nuclear Research) (2026) | TGRS Research Map | TGRS