ProactiveAudioBench: Evaluating Condition-Driven Responses in Audio Streams
Audio assistants are moving from turn-based question answering toward continuous listening, where they must interpret incoming audio while deciding whether to respond or remain silent. Existing proactive benchmarks emphasize video and audiovisual interaction, while audio-centric evaluations offer limited systematic coverage of listening conditions across audio domains. We introduce ProactiveAudioBench, a benchmark for condition-driven responses across sound events, speech, and music. Four tasks—Detection, Tracking, Counting, and Matching—cover first-occurrence, repeated, count-conditioned, and reference-based notifications. Natural-language conditions and optional acoustic references are paired with annotated response opportunities and reference replies, supporting both single-response and multiple-response requests. The framework accommodates simulated and native streaming and defines joint time-and-content Coverage alongside negative-request False Alarm Rate. An initial offline study evaluates seven open-source audio models across 56 subsets, yielding 27,120 prediction records. Using first-trigger diagnostics, the highest reported timing F1 at a ±3s tolerance is 0.292. Baichuan-Omni-1.5 and Qwen Thinking have negative-request false-trigger rates of 0.581 and 0.861 on their respective request sets, whereas the other five models have response-presence F1 no greater than 0.113. These subset-level results reveal difficulty balancing timely notifications with appropriate silence. ProactiveAudioBench provides a testbed for connecting audio understanding with reliable execution of persistent listening instructions.
Authors
- Ling Wang (ORCID: https://orcid.org/0000-0001-8964-6454)
- Yuxuan Wang (ORCID: https://orcid.org/0009-0002-5716-6776)
- Qi Chen
- Meng Gao
- Yinsong Yan
- Wenxiang Guo
- Yunfei Chu
- Jin Xu
- Yifan Yang
- Yujiu Yang
- Hui Wang
- Zihan Liu
- Wen Huang
- Xize Cheng
Institutions
- Hong Kong Polytechnic University (HK)
- Johns Hopkins University (US)
- Shanghai Jiao Tong University (CN)
- Nankai University (CN)
- Alibaba Group (United States) (US)
- Zhejiang University (CN)
- Tsinghua University (CN)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-21
- DOI
- https://doi.org/10.5281/zenodo.22871964
- Primary Topic
- Speech and Audio Processing
- Type
- preprint