A Survey of Visual Question Answering for Embodied Robots: Tasks, Methods, and Future Directions

Visual Question Answering (VQA) in embodied settings draws on computer vision, natural language processing, and robotics. Unlike traditional VQA, Embodied Question Answering (EQA) may require an agent to acquire and retain evidence from a 3D environment before answering a natural language query. This semi-systematic survey reviews a core corpus of 72 papers selected from 210 candidates and supplements it with a separate targeted qualitative update comprising 28 additional records identified through July 2026. We use PMRA (Perception–Memory–Reasoning–Action) as an author-developed analytical framework rather than a new robot-control architecture, together with a three-level taxonomy of task formulations, method architectures, and capability dimensions. An audit of the 20 tabulated datasets, with counts recomputed from the accompanying extraction sheet, finds that five permit or require active exploration and only one requires physical object interaction. We also define four architecture-based stages without using publication year as an assignment rule. The qualitative synthesis identifies examples of explicit memory interfaces and question-conditioned information-acquisition mechanisms, but it does not infer their prevalence, temporal growth, or relative importance from publication frequencies. Finally, we discuss four recurring limitations of foundation-model approaches and nine research directions toward 2030.

Authors

Institutions

Publication Details

Journal
AI
Published
2026-09-15
DOI
https://doi.org/10.3390/ai7090367
Primary Topic
Multimodal Machine Learning Applications
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

A Survey of Visual Question Answering for Embodied Robots: Tasks, Methods, and Future Directions

Yinlong Liu, Weihan Shi
AI
Multimodal Machine Learning Applications
article

A Survey of Visual Question Answering for Embodied Robots: Tasks, Methods, and Future Directions

Yinlong Liu, Weihan Shi
article en

Abstract

Visual Question Answering (VQA) in embodied settings draws on computer vision, natural language processing, and robotics. Unlike traditional VQA, Embodied Question Answering (EQA) may require an agent to acquire and retain evidence from a 3D environment before answering a natural language query. This semi-systematic survey reviews a core corpus of 72 papers selected from 210 candidates and supplements it with a separate targeted qualitative update comprising 28 additional records identified through July 2026. We use PMRA (Perception–Memory–Reasoning–Action) as an author-developed analytical framework rather than a new robot-control architecture, together with a three-level taxonomy of task formulations, method architectures, and capability dimensions. An audit of the 20 tabulated datasets, with counts recomputed from the accompanying extraction sheet, finds that five permit or require active exploration and only one requires physical object interaction. We also define four architecture-based stages without using publication year as an assignment rule. The qualitative synthesis identifies examples of explicit memory interfaces and question-conditioned information-acquisition mechanisms, but it does not infer their prevalence, temporal growth, or relative importance from publication frequencies. Finally, we discuss four recurring limitations of foundation-model approaches and nine research directions toward 2030.

AIVol. 7(9)
Minzu University of China (CN), City University of Macau (MO)
Quality Education
Openalex Percentile: Top 13%
Multimodal Machine Learning Applications
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

A Survey of Visual Question Answering for Embodied Robots: Tasks, Methods, and Future Directions — Yinlong Liu, Weihan Shi · AI (2026) | TGRS Research Map | TGRS