Cross-Lingual Real-Time Speech Interaction for Embodied Intelligence: Architecture, Challenges, and Prospects
Embodied intelligence is rapidly evolving toward natural human–machine interaction, where real-time cross-lingual speech processing is essential for robots to understand and execute commands across languages. This study reviews research progress from January 2010 to November 1, 2025, using CiteSpace to analyze knowledge structure, hotspots, and trends. The year 2010 was selected as the start date because it marks the early emergence of foundational theoretical discussions on embodied intelligence internationally, providing a comprehensive baseline to observe the entire evolutionary trajectory of the field, even though substantive domestic research emerged later around 2015. Results show a sharp rise in research interest after 2023, driven by multilingual interaction and robotics applications. Current hotspots include multimodal representation learning, low-latency speech-to-text models, cross-lingual semantic alignment, and embodied control. The field has progressed through three stages: conceptual foundations, real-time neural integration, and large-model-driven general embodiment. We summarize the technical framework spanning speech recognition, semantic grounding, and action execution, and identify challenges such as latency reduction, noise robustness, and cross-lingual knowledge transfer. Future work should focus on unified speech–action representation, scalable low-latency architectures, adaptive learning, and ethical governance. This study outlines development trends and prospects for cross-lingual real-time speech interaction in embodied intelligence.
Authors
- yicheng li (ORCID: https://orcid.org/0000-0003-1186-363X)
- Guanzhe Jiao (ORCID: https://orcid.org/0000-0003-2424-116X)
- Shuai Wang (ORCID: https://orcid.org/0000-0002-7884-8920)
Institutions
- Shandong Management University (CN)
- Shandong Women’s University (CN)
- Southwest Jiaotong University (CN)
Publication Details
- Journal
- Big Data
- Published
- 2026-10-07
- DOI
- https://doi.org/10.1177/2167647x261487329
- Primary Topic
- Speech Recognition and Synthesis
- Type
- article
- Field-Weighted Citation Impact
- 0.00