keywords:
spatial cognition
cognitive neuroscience
fmri
artificial intelligence
knowledge representation
Multimodal models excel in tasks requiring semantic integra- tion of language and vision but struggle with spatial cognition. Using a visual perspective-taking task inspired by cognitive science, we find these models fail when the image and ref- erence view differ, reflecting spatial cognition comparable to a two-year-old child. To explore these disparities further, we analyze internal representations using a human action fMRI dataset and voxelwise encoding models, revealing key differ- ences between AI and human spatial encoding. This work pro- vides new benchmarks and insights into bridging artificial and biological cognition.
