Development of a Deep Learning Model for Audio Question Answering (AQA)
This paper presents a high-efficiency multimodal framework designed for Audio Question Answering (AQA). To overcome standard CPU and disk I/O bottlenecks, the architecture implements a zero-I/O data pipeline that pre-processes and caches 10,686 samples directly into the 48 GB VRAM of an NVIDIA RTX 6000 Ada Generation GPU during initialization. The model combines a 2D Residual Network (ResNet) audio encoder with an unfrozen 6-layer BERT-base text encoder using a Bilinear Gated Fusion mechanism. Trained across 15 epochs with EMA weight stabilization, the system achieves an extraction throughput of 128-131 files per second and epoch durations of 30-40 seconds, establishing a robust baseline for multimodal acoustic reasoning tasks.
Authors
- Vikash Kumar Singh (ORCID: https://orcid.org/0000-0002-8747-1627)
- Kalpana Sharma (ORCID: https://orcid.org/0000-0001-5083-2428)
- Pritam Barua
- Samarendra Nath Sur Sur
- Mridul Mayank Jha
Institutions
- Sikkim Manipal University (IN)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-05
- DOI
- https://doi.org/10.5281/zenodo.22341166
- Primary Topic
- Music and Audio Processing
- Type
- article
- Field-Weighted Citation Impact
- 0.00