Deep-learning-based fMRI decoding of real-world size for hand-held objects
Real-world object size is a fundamental dimension of visual cognition, supporting effective interaction with the environment and object manipulation. However, neural mechanisms underlying size representation have largely been inferred from extreme size comparisons, leaving the neural representation of subtle size differences within a manipulable, “hand-scale” range poorly understood. Here, we applied a three-dimensional deep neural network (3D DNN) to decode real-world size from whole-brain fMRI data ( N = 50) using objects that all fell within a graspable range. The 3D DNN successfully decoded subtle size differences, achieving predictive accuracies comparable to those obtained with multivariate pattern analysis. Importantly, Guided Gradient-weighted Class Activation Mapping (Guided Grad-CAM) revealed that voxel patterns contributing to the classifier’s small- versus large-object predictions were not confined to the ventral occipito-temporal cortex but extended to distributed regions. Notably, these regions spatially overlapped with the specific brain areas previously implicated in size-perception distortions following brain damage. Our findings suggest that the subtle variations in object size may be represented through a non-linear, distributed network that extends beyond the traditional visual hierarchy. Specifically, this system may support the integration of visual properties with semantic scaling and the multimodal convergence of vision, space, and memory.
Authors
- Kosuke Miyoshi (ORCID: https://orcid.org/0000-0002-8237-1395)
- Cui Lang (ORCID: https://orcid.org/0009-0003-2168-5096)
- Masaki Takeda
Institutions
- Kochi University of Technology (JP)
Publication Details
- Journal
- Scientific Reports
- Published
- 2026-10-07
- DOI
- https://doi.org/10.1038/s41598-026-74629-7
- Primary Topic
- Visual perception and processing mechanisms
- Type
- article
- Field-Weighted Citation Impact
- 0.00