A Survey of Visual Question Answering for Embodied Robots: Tasks, Methods, and Future Directions
Visual Question Answering (VQA) in embodied settings draws on computer vision, natural language processing, and robotics. Unlike traditional VQA, Embodied Question Answering (EQA) may require an agent to acquire and retain evidence from a 3D environment before answering a natural language query. This semi-systematic survey reviews a core corpus of 72 papers selected from 210 candidates and supplements it with a separate targeted qualitative update comprising 28 additional records identified through July 2026. We use PMRA (Perception–Memory–Reasoning–Action) as an author-developed analytical framework rather than a new robot-control architecture, together with a three-level taxonomy of task formulations, method architectures, and capability dimensions. An audit of the 20 tabulated datasets, with counts recomputed from the accompanying extraction sheet, finds that five permit or require active exploration and only one requires physical object interaction. We also define four architecture-based stages without using publication year as an assignment rule. The qualitative synthesis identifies examples of explicit memory interfaces and question-conditioned information-acquisition mechanisms, but it does not infer their prevalence, temporal growth, or relative importance from publication frequencies. Finally, we discuss four recurring limitations of foundation-model approaches and nine research directions toward 2030.
Authors
- Yinlong Liu (ORCID: https://orcid.org/0000-0002-6468-8233)
- Weihan Shi (ORCID: https://orcid.org/0009-0009-0593-3639)
Institutions
- Minzu University of China (CN)
- City University of Macau (MO)
Publication Details
- Journal
- AI
- Published
- 2026-09-15
- DOI
- https://doi.org/10.3390/ai7090367
- Primary Topic
- Multimodal Machine Learning Applications
- Type
- article
- Field-Weighted Citation Impact
- 0.00