Latest Research in Conversational Speech Processing

360 research papers · 0.0 average citations · 2026 median publication year

Top Research Topics in Conversational Speech Processing

Highest-Cited Papers

  1. DOTA-ME-CS: daily oriented text audio-Mandarin English-Code switching dataset (1 citations)
  2. Towards inclusive voice biometrics: Dysarthria-discriminative embeddings for ASV system
  3. Speech intelligibility comparison of standalone and two-stage deep learning architectures for behind-the-ear-to-binaural enhancement
  4. A2SSC: An Agent-based Adaptive Semantic Speech Communication System
  5. DESED and DataSED precomputed caches for Cover First, Disagree Softly: Rethinking Mismatch-First Active Learning for Frame-Level Audio Classification
  6. Enabling automatic transcription of child-centered audio recordings from real-world environments
  7. Model-Agnostic and Language-Agnostic Voice Pipeline Improvement for the Agriculture Domain
  8. Foreground Voice Activity Detection: Learning Speaker Selectivity from Supervision
  9. Personalized Keyword Spotting for User-Defined Keywords Leveraging Text-Independent Speaker Verification
  10. CircleMatch: Prototype Matching with Circular Temporal Statistics for Tiny Keyword Spotting
  11. Multimodal Conversational Context for LLM-Based ASR: Data Construction, Training, and Benchmark
  12. Can Large Audio Language Models Ignore Multilingual Distractors? An Evaluation of Their Selective Auditory Attention Capabilities
  13. Alignment-Path Distillation from Non-streaming ASR-LLMs for Streaming Speech Recognition
  14. Evaluating Bias in Phoneme-Based Automatic Speech Recognition Systems: An Analysis of IPA Transcription Models
  15. Scaling Audio Models Efficiently: Joint Optimization of Scale, Resolution, Adaptation, Precision, and Sparsity
  16. Phoneme-guided TTS augmentation for ASR: A unified pipeline and multilingual evaluation
  17. Soft Posterior Speaker Injection for Multi-Talker Speech Recognition
  18. Design of the IBM Granite 5.0 TurboCTC ASR Model
  19. Multi-Dimensional Prosody Judgment For Live Streaming Speech Synthesis
  20. Beyond the Stability--Plasticity Frontier in Streaming Target Speaker Extraction

Sub-Regions

Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
L2 Region - - 2026 Sep Q3

Conversational Speech Processing

360 papers

Top Topics (10)

Audio and Speech Processing110
Sound97
Computation and Language74
Speech Recognition and Synthesis24
Machine Learning11
Artificial Intelligence7
Natural Language Processing Techniques5
Emotion and Mood Recognition5
Computer Vision and Pattern Recognition5
Phonetics and Phonology Research5

Top Publications (20)

1.DOTA-ME-CS: daily oriented text audio-Mandarin English-Code switching dataset1c2.Towards inclusive voice biometrics: Dysarthria-discriminative embeddings for ASV system3.Speech intelligibility comparison of standalone and two-stage deep learning architectures for behind-the-ear-to-binaural enhancement4.A2SSC: An Agent-based Adaptive Semantic Speech Communication System5.DESED and DataSED precomputed caches for Cover First, Disagree Softly: Rethinking Mismatch-First Active Learning for Frame-Level Audio Classification6.Enabling automatic transcription of child-centered audio recordings from real-world environments7.Model-Agnostic and Language-Agnostic Voice Pipeline Improvement for the Agriculture Domain8.Foreground Voice Activity Detection: Learning Speaker Selectivity from Supervision9.Personalized Keyword Spotting for User-Defined Keywords Leveraging Text-Independent Speaker Verification10.CircleMatch: Prototype Matching with Circular Temporal Statistics for Tiny Keyword Spotting11.Multimodal Conversational Context for LLM-Based ASR: Data Construction, Training, and Benchmark12.Can Large Audio Language Models Ignore Multilingual Distractors? An Evaluation of Their Selective Auditory Attention Capabilities13.Alignment-Path Distillation from Non-streaming ASR-LLMs for Streaming Speech Recognition14.Evaluating Bias in Phoneme-Based Automatic Speech Recognition Systems: An Analysis of IPA Transcription Models15.Scaling Audio Models Efficiently: Joint Optimization of Scale, Resolution, Adaptation, Precision, and Sparsity16.Phoneme-guided TTS augmentation for ASR: A unified pipeline and multilingual evaluation17.Soft Posterior Speaker Injection for Multi-Talker Speech Recognition18.Design of the IBM Granite 5.0 TurboCTC ASR Model19.Multi-Dimensional Prosody Judgment For Live Streaming Speech Synthesis20.Beyond the Stability--Plasticity Frontier in Streaming Target Speaker Extraction

Sub-Regions (6)

AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.