SoundTrace: Integrating Temporal Context and Episodic Memory for Real-Time Sound Recognition
Environmental sound recognition systems have become increasingly capable, yet they often operate in context-free modes that ignore temporal continuity, environmental patterns, and user feedback. We present SoundTrace, a real-time sound recognition system that integrates temporal context and episodic memory to support adaptive, interpretable inference in everyday environments. SoundTrace augments a neural audio classifier with lightweight memory structures that store symbolic event traces, estimate scene-time and short-range sequence priors, and update those priors through user feedback. During inference, these memory-derived priors are retrieved and fused with model predictions to stabilize labels and provide interpretable reasoning. In a controlled evaluation, we show that contextual inference improves accuracy, reduces label volatility, and enhances robustness under ambiguous conditions. We also report findings from an eight-week in-home deployment with 14 deaf and hard-of-hearing participants, revealing how context-aware feedback and explanations shape users' trust, understanding, and correction strategies. Our results demonstrate the viability of context-integrated sound recognition in everyday environments.
Authors
- Dhruv Jain (ORCID: https://orcid.org/0000-0001-6176-968X)
- Jason Miller (ORCID: https://orcid.org/0009-0001-1623-8553)
Institutions
- University of Michigan (US)
Publication Details
- Journal
- Proceedings of the ACM on Interactive Mobile Wearable and Ubiquitous Technologies
- Published
- 2026-09-30
- DOI
- https://doi.org/10.1145/3831645
- Primary Topic
- Music and Audio Processing
- Type
- article
- Field-Weighted Citation Impact
- 0.00