Opening the Window of the Mind: Do Large Language Models Possess Human-Like Visual Attention?

When people read a spatial description, their eyes move through empty space as if the described scene were present, a signature of imagery-driven attention documented by Spivey and Geng. Language models have no eyes, but they are routinely credited with spatial understanding on the basis of verbal answers, and such answers are ambiguous: a model can name the right direction by echoing a directional word the passage happens to contain. We ask whether responses to spatial descriptions are governed by the configuration described or by the vocabulary used to describe it. Adapting the Spivey and Geng materials, we hold each scene’s layout fixed and vary its linguistic support across four versions: one naming the direction, one conveying only the configuration, one reordering the same sentences, and one inserting a directional word that contradicts the layout. Five models from the Qwen and GLM families were probed with a forced-choice direction report, an unconstrained narration, a two-dimensional reconstruction, and a plain relational question, together with control conditions that vary the position and register of the inserted word, shuffle the object list in the prompt, remove the imagery framing, and offer an explicit indeterminate option. Analysis treats model-by-scene cells as units, with cluster-bootstrap intervals and mixed-effects models, across 14,900 responses. Passages that convey a layout without naming a direction are answered near ceiling, reconstructed and narrated in the correct object order, and unaffected by reordering, by shuffling the prompt’s object list, or by removing the imagery framing: directional vocabulary is not what these responses track. Three apparent limits on that competence proved to be artifacts of measurement. A misleading word is resisted on three quarters of trials when it precedes the layout and on none when it follows it, so the manipulation indexes recency rather than representational strength; a downward bias on axis-free scenes disappears once declining is permitted; and narrations judged unfaithful by direction coding are faithful when scored by object order. Layout recovery is robust, and spatial evaluations that vary a single response format risk measuring their instrument.

Authors

Institutions

Publication Details

Journal
Journal of Intelligence
Published
2026-10-01
DOI
https://doi.org/10.3390/jintelligence14100234
Primary Topic
Categorization, perception, and language
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Opening the Window of the Mind: Do Large Language Models Possess Human-Like Visual Attention?

Huanghao Feng, Jian Wang, Houji Jin, Ziyi Wang et al.
Journal of Intelligence
Categorization, perception, and language
article

Opening the Window of the Mind: Do Large Language Models Possess Human-Like Visual Attention?

Huanghao Feng, Jian Wang, Houji Jin, Ziyi Wang, Jiayi Zhang, Zhicheng Cai, Ganyu Gui, Zhijie Tang
article en

Abstract

When people read a spatial description, their eyes move through empty space as if the described scene were present, a signature of imagery-driven attention documented by Spivey and Geng. Language models have no eyes, but they are routinely credited with spatial understanding on the basis of verbal answers, and such answers are ambiguous: a model can name the right direction by echoing a directional word the passage happens to contain. We ask whether responses to spatial descriptions are governed by the configuration described or by the vocabulary used to describe it. Adapting the Spivey and Geng materials, we hold each scene’s layout fixed and vary its linguistic support across four versions: one naming the direction, one conveying only the configuration, one reordering the same sentences, and one inserting a directional word that contradicts the layout. Five models from the Qwen and GLM families were probed with a forced-choice direction report, an unconstrained narration, a two-dimensional reconstruction, and a plain relational question, together with control conditions that vary the position and register of the inserted word, shuffle the object list in the prompt, remove the imagery framing, and offer an explicit indeterminate option. Analysis treats model-by-scene cells as units, with cluster-bootstrap intervals and mixed-effects models, across 14,900 responses. Passages that convey a layout without naming a direction are answered near ceiling, reconstructed and narrated in the correct object order, and unaffected by reordering, by shuffling the prompt’s object list, or by removing the imagery framing: directional vocabulary is not what these responses track. Three apparent limits on that competence proved to be artifacts of measurement. A misleading word is resisted on three quarters of trials when it precedes the layout and on none when it follows it, so the manipulation indexes recency rather than representational strength; a downward bias on axis-free scenes disappears once declining is permitted; and narrations judged unfaithful by direction coding are faithful when scored by object order. Layout recovery is robust, and spatial evaluations that vary a single response format risk measuring their instrument.

Journal of IntelligenceVol. 14(10)
Hong Kong Polytechnic University (HK), Suzhou University of Technology (CN), Suzhou University of Science and Technology (CN)
Reduced inequalities
Openalex Percentile: Top 8%
Categorization, perception, and language
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.