CogSci 2025

•

August 02, 2025

•

San Francisco, United States

keywords:

spatial cognition

cognitive neuroscience

fmri

artificial intelligence

knowledge representation

Multimodal models excel in tasks requiring semantic integra- tion of language and vision but struggle with spatial cognition. Using a visual perspective-taking task inspired by cognitive science, we find these models fail when the image and ref- erence view differ, reflecting spatial cognition comparable to a two-year-old child. To explore these disparities further, we analyze internal representations using a human action fMRI dataset and voxelwise encoding models, revealing key differ- ences between AI and human spatial encoding. This work pro- vides new benchmarks and insights into bridging artificial and biological cognition.

Downloads

PaperTranscript English (automatic)

Next from CogSci 2025

Do our theories of moral progress predict whether we vote?  Evidence from the 2024 US election
poster

Do our theories of moral progress predict whether we vote? Evidence from the 2024 US election

CogSci 2025

Casey Lewry
Tania Lombrozo
Casey Lewry and 1 other author

02 August 2025

Similar lecture

Spatial Representation of Large Language Models in 2D Scene
workshop paper

Spatial Representation of Large Language Models in 2D Scene

ACL 2025

Wenya Wu

31 July 2025