CogSci 2025

•

August 01, 2025

•

San Francisco, United States

keywords:

human factors

pattern recognition

eye tracking

artificial intelligence

vision

Inspired by human visual attention, deep neural networks have widely adopted attention mechanisms to learn locally discriminative attributes for challenging visual classification tasks. However, existing approaches primarily emphasize the representation of such features while neglecting precise localization, which often leads to misclassification caused by shortcut biases. This limitation becomes more pronounced when models are evaluated on transfer or out-of-distribution datasets. In contrast, humans leverage prior object knowledge to quickly localize and compare fine-grained attributes, a capability especially crucial in complex and high-variance classification scenarios. We introduce Gaze-CIFAR-10, a human gaze time-series dataset, along with a dual-sequence gaze encoder that models the precise sequential localization of human attention on distinct local attributes. In parallel, a Vision Transformer (ViT) is employed to learn the sequential representation of image content. Through cross-modal fusion, our framework integrates human gaze priors with machine-derived visual sequences, effectively correcting inaccurate localization in image feature representations.

Downloads

PaperTranscript English (automatic)

Next from CogSci 2025

DHRec: A Debiased Hyperbolic Recommendation Model
poster

DHRec: A Debiased Hyperbolic Recommendation Model

CogSci 2025

Xianglong Li
+3
Mengmeng Li and 5 other authors

01 August 2025

Similar lecture

CLIP-Gaze: Towards General Gaze Estimation via Visual-Linguistic Model
technical paper

CLIP-Gaze: Towards General Gaze Estimation via Visual-Linguistic Model

AAAI 2024

Pengwei Yin
+1
Pengwei Yin and 3 other authors

24 February 2024