ProactiveAudioBench: Evaluating Condition-Driven Responses in Audio Streams

Audio assistants are moving from turn-based question answering toward continuous listening, where they must interpret incoming audio while deciding whether to respond or remain silent. Existing proactive benchmarks emphasize video and audiovisual interaction, while audio-centric evaluations offer limited systematic coverage of listening conditions across audio domains. We introduce ProactiveAudioBench, a benchmark for condition-driven responses across sound events, speech, and music. Four tasks—Detection, Tracking, Counting, and Matching—cover first-occurrence, repeated, count-conditioned, and reference-based notifications. Natural-language conditions and optional acoustic references are paired with annotated response opportunities and reference replies, supporting both single-response and multiple-response requests. The framework accommodates simulated and native streaming and defines joint time-and-content Coverage alongside negative-request False Alarm Rate. An initial offline study evaluates seven open-source audio models across 56 subsets, yielding 27,120 prediction records. Using first-trigger diagnostics, the highest reported timing F1 at a ±3s tolerance is 0.292. Baichuan-Omni-1.5 and Qwen Thinking have negative-request false-trigger rates of 0.581 and 0.861 on their respective request sets, whereas the other five models have response-presence F1 no greater than 0.113. These subset-level results reveal difficulty balancing timely notifications with appropriate silence. ProactiveAudioBench provides a testbed for connecting audio understanding with reliable execution of persistent listening instructions.

Authors

Institutions

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-21
DOI
https://doi.org/10.5281/zenodo.22871964
Primary Topic
Speech and Audio Processing
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
preprint

ProactiveAudioBench: Evaluating Condition-Driven Responses in Audio Streams

Ling Wang, Yuxuan Wang, Qi Chen, Meng Gao et al.
Zenodo (CERN European Organization for Nuclear Research)
Speech and Audio Processing
preprint

ProactiveAudioBench: Evaluating Condition-Driven Responses in Audio Streams

Ling Wang, Yuxuan Wang, Qi Chen, Meng Gao, Yinsong Yan, Wenxiang Guo, Yunfei Chu, Jin Xu, Yifan Yang, Yujiu Yang, Hui Wang, Zihan Liu, Wen Huang, Xize Cheng
preprint en

Abstract

Audio assistants are moving from turn-based question answering toward continuous listening, where they must interpret incoming audio while deciding whether to respond or remain silent. Existing proactive benchmarks emphasize video and audiovisual interaction, while audio-centric evaluations offer limited systematic coverage of listening conditions across audio domains. We introduce ProactiveAudioBench, a benchmark for condition-driven responses across sound events, speech, and music. Four tasks—Detection, Tracking, Counting, and Matching—cover first-occurrence, repeated, count-conditioned, and reference-based notifications. Natural-language conditions and optional acoustic references are paired with annotated response opportunities and reference replies, supporting both single-response and multiple-response requests. The framework accommodates simulated and native streaming and defines joint time-and-content Coverage alongside negative-request False Alarm Rate. An initial offline study evaluates seven open-source audio models across 56 subsets, yielding 27,120 prediction records. Using first-trigger diagnostics, the highest reported timing F1 at a ±3s tolerance is 0.292. Baichuan-Omni-1.5 and Qwen Thinking have negative-request false-trigger rates of 0.581 and 0.861 on their respective request sets, whereas the other five models have response-presence F1 no greater than 0.113. These subset-level results reveal difficulty balancing timely notifications with appropriate silence. ProactiveAudioBench provides a testbed for connecting audio understanding with reliable execution of persistent listening instructions.

Zenodo (CERN European Organization for Nuclear Research)
Hong Kong Polytechnic University (HK), Johns Hopkins University (US), Shanghai Jiao Tong University (CN), Nankai University (CN), Alibaba Group (United States) (US), Zhejiang University (CN), Tsinghua University (CN)
Speech and Audio Processing
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.