Latest Research in Audio and Speech Processing
35 research papers · 2026 median publication year
Top Research Topics in Audio and Speech Processing
- Sound — 14 papers
- Audio and Speech Processing — 7 papers
- Music and Audio Processing — 2 papers
- Signal Processing — 2 papers
- Computation and Language — 2 papers
- EEG and Brain-Computer Interfaces — 2 papers
- Machine Fault Diagnosis Techniques — 1 papers
- Hearing Loss and Rehabilitation — 1 papers
- Computer Vision and Pattern Recognition — 1 papers
- Machine Learning — 1 papers
Highest-Cited Papers
- DESED and DataSED precomputed caches for Cover First, Disagree Softly: Rethinking Mismatch-First Active Learning for Frame-Level Audio Classification
- Arti-JEPA: Adapting Video World Model to Real-Time MRI of the Vocal Tract for Speech-Production Analysis
- PitchFlower: A flow-based neural audio codec with pitch controllability
- Task-oriented neural FOA encoding for SELD from irregular microphone arrays
- SyncVoice: Simple and Effective Automatic Video Dubbing with Vision-Augmented TTS
- Audio-Visual Turn-taking Prediction in Cocktail Party Scenarios
- Tracing the Origins: Legacy Codec Identification in Neural Audio Transcoding
- Multiscale Gaussian-Mixture Modeling for HMM Post-Processing in Selective Auditory Attention Decoding
- OLAC: An Overlapped Lossless Audio Codec in the Time-Domain with MDCT Compatibility
- ECHOv2: A Frequency-Structured Pre-trained Acoustic Representation Model with Cross-Band Modeling for Machine Anomalous Sound Detection
- LoSATok: Low-Dimensional Semantic-Acoustic Tokenizer for Cross-Domain Audio Understanding and Generation
- Adaptive Perturbation Selection for Contrastive Audio Decoding
- The Eloquence submission for Task 2 of the Interspeech 2026 MLC-SLM challenge
- Domain-Adaptive Audio Large Language Model for Acoustic Fault Diagnosis and Semantic Description of Coal Mine Equipment
- The Semantic Bottleneck: Leveraging Semantic Representations for Non-Invasive Speech Decoding
- Learned Continuous Synthesis of Quadratic Difference Tone Spectra
- One-Stage Multi-Task Instruction-Guided 3D Spatial Audio Editing
- Do EEG Foundation Models Transfer to Speech? A Benchmark on Overt and Imagined Speech Decoding
- Iterative Audio Separation with Mixture Consistency via MIMO Model Extension
- Efficient and Microphone-Fault-Tolerant 3D Sound Source Localization