Latest Research in Speech Recognition and Synthesis
54 research papers · 0.0 average citations · 2026 median publication year
Top Research Topics in Speech Recognition and Synthesis
- Audio and Speech Processing — 16 papers
- Computation and Language — 13 papers
- Sound — 12 papers
- Speech Recognition and Synthesis — 7 papers
- Emotion and Mood Recognition — 2 papers
- Artificial Intelligence — 1 papers
- Phonetics and Phonology Research — 1 papers
- Machine Learning — 1 papers
- Computer Vision and Pattern Recognition — 1 papers
Highest-Cited Papers
- DOTA-ME-CS: daily oriented text audio-Mandarin English-Code switching dataset (1 citations)
- A2SSC: An Agent-based Adaptive Semantic Speech Communication System
- Personalized Keyword Spotting for User-Defined Keywords Leveraging Text-Independent Speaker Verification
- Multimodal Conversational Context for LLM-Based ASR: Data Construction, Training, and Benchmark
- Self-Distilled Pronunciation and Accent Control for Neural Text-to-Speech
- VoiceTrace: A Benchmark and Retrieval Framework for Who-Said-What Speech Retrieval
- A frontend-backend architecture for tool calls in full-duplex speech models
- Causal Analysis and Mitigation of Spurious Onsets in Full-Duplex Speech LLMs
- SPEAKER ACCENT RECOGNITION USING MACHINE LEARNING
- SPEAKER ACCENT RECOGNITION USING MACHINE LEARNING
- Empirical mode decomposition based CNN prediction model for speech emotion classification using Mel scale spectrograms
- Taming Long-form Text-to-Speech
- CLASH: Counterfactual Auditing of Lexical and Prosodic Reliance in Spoken Sarcasm Detection
- SITA: Learning Speaker-Invariant and Tone-Aware Speech Representations for Low-Resource Tonal Languages
- Enhancing Law-Enforcement Audio Transcription: A LoRA-Based Adaptation of Whisper for BWC Footage
- TargetSEC: Plug-and-Play In-the-Wild Speech Emotion Conversion via Arousal-Conditioned Latent Style Diffusion
- Enabling Streaming User Transcription in Full-Duplex Speech-to-Speech Models
- Word Timestamps and Speaker Attribution with a Non-Autoregressive LLM
- Tone on a Budget: A Reference-Free Metric for Lexical Tone in Massively Multilingual Text-to-Speech
- Neyshekar: An Open Persian Read-Speech Corpus for Automatic Speech Recognition