Latest Research in Audio and Speech Processing

35 research papers · 2026 median publication year

Top Research Topics in Audio and Speech Processing

Highest-Cited Papers

  1. DESED and DataSED precomputed caches for Cover First, Disagree Softly: Rethinking Mismatch-First Active Learning for Frame-Level Audio Classification
  2. Arti-JEPA: Adapting Video World Model to Real-Time MRI of the Vocal Tract for Speech-Production Analysis
  3. PitchFlower: A flow-based neural audio codec with pitch controllability
  4. Task-oriented neural FOA encoding for SELD from irregular microphone arrays
  5. SyncVoice: Simple and Effective Automatic Video Dubbing with Vision-Augmented TTS
  6. Audio-Visual Turn-taking Prediction in Cocktail Party Scenarios
  7. Tracing the Origins: Legacy Codec Identification in Neural Audio Transcoding
  8. Multiscale Gaussian-Mixture Modeling for HMM Post-Processing in Selective Auditory Attention Decoding
  9. OLAC: An Overlapped Lossless Audio Codec in the Time-Domain with MDCT Compatibility
  10. ECHOv2: A Frequency-Structured Pre-trained Acoustic Representation Model with Cross-Band Modeling for Machine Anomalous Sound Detection
  11. LoSATok: Low-Dimensional Semantic-Acoustic Tokenizer for Cross-Domain Audio Understanding and Generation
  12. Adaptive Perturbation Selection for Contrastive Audio Decoding
  13. The Eloquence submission for Task 2 of the Interspeech 2026 MLC-SLM challenge
  14. Domain-Adaptive Audio Large Language Model for Acoustic Fault Diagnosis and Semantic Description of Coal Mine Equipment
  15. The Semantic Bottleneck: Leveraging Semantic Representations for Non-Invasive Speech Decoding
  16. Learned Continuous Synthesis of Quadratic Difference Tone Spectra
  17. One-Stage Multi-Task Instruction-Guided 3D Spatial Audio Editing
  18. Do EEG Foundation Models Transfer to Speech? A Benchmark on Overt and Imagined Speech Decoding
  19. Iterative Audio Separation with Mixture Consistency via MIMO Model Extension
  20. Efficient and Microphone-Fault-Tolerant 3D Sound Source Localization
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
L3 Region - - 2026 Sep Q3

Audio and Speech Processing

35 papers

Top Topics (10)

Sound14
Audio and Speech Processing7
Music and Audio Processing2
Signal Processing2
Computation and Language2
EEG and Brain-Computer Interfaces2
Machine Fault Diagnosis Techniques1
Hearing Loss and Rehabilitation1
Computer Vision and Pattern Recognition1
Machine Learning1

Top Publications (20)

1.DESED and DataSED precomputed caches for Cover First, Disagree Softly: Rethinking Mismatch-First Active Learning for Frame-Level Audio Classification2.Arti-JEPA: Adapting Video World Model to Real-Time MRI of the Vocal Tract for Speech-Production Analysis3.PitchFlower: A flow-based neural audio codec with pitch controllability4.Task-oriented neural FOA encoding for SELD from irregular microphone arrays5.SyncVoice: Simple and Effective Automatic Video Dubbing with Vision-Augmented TTS6.Audio-Visual Turn-taking Prediction in Cocktail Party Scenarios7.Tracing the Origins: Legacy Codec Identification in Neural Audio Transcoding8.Multiscale Gaussian-Mixture Modeling for HMM Post-Processing in Selective Auditory Attention Decoding9.OLAC: An Overlapped Lossless Audio Codec in the Time-Domain with MDCT Compatibility10.ECHOv2: A Frequency-Structured Pre-trained Acoustic Representation Model with Cross-Band Modeling for Machine Anomalous Sound Detection11.LoSATok: Low-Dimensional Semantic-Acoustic Tokenizer for Cross-Domain Audio Understanding and Generation12.Adaptive Perturbation Selection for Contrastive Audio Decoding13.The Eloquence submission for Task 2 of the Interspeech 2026 MLC-SLM challenge14.Domain-Adaptive Audio Large Language Model for Acoustic Fault Diagnosis and Semantic Description of Coal Mine Equipment15.The Semantic Bottleneck: Leveraging Semantic Representations for Non-Invasive Speech Decoding16.Learned Continuous Synthesis of Quadratic Difference Tone Spectra17.One-Stage Multi-Task Instruction-Guided 3D Spatial Audio Editing18.Do EEG Foundation Models Transfer to Speech? A Benchmark on Overt and Imagined Speech Decoding19.Iterative Audio Separation with Mixture Consistency via MIMO Model Extension20.Efficient and Microphone-Fault-Tolerant 3D Sound Source Localization
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.