CogSci 2025

•

August 01, 2025

•

San Francisco, United States

keywords:

predictive processing

language comprehension

language production

artificial intelligence

pragmatics

Contemporary transformer models have achieved human-like performance on many text-based tasks. However, real-world communication requires the integration of language with non-linguistic context (e.g., visual, social, etc.). Here, we study such information integration in three multimodal transformer models. We test these models’ pragmatic capabilities regarding referring expressions: when an object set contains two exemplars from the same category that differ in size, unambiguously referring to one of them requires a size adjective (e.g., the big hammer); the adjective is unnecessary if only one exemplar from the category is present. We evaluate these inferences when models process text-image inputs (via their surprisal for infelicitous vs. felicitous adjective use) and when they generate open-ended descriptions of images given text prompts. We find evidence for pragmatic integration of visual and linguistic context in all models. However, these inferences remain sensitive to the in-context statistics of visual inputs, unlike pragmatic inference in humans.

Downloads

Paper

Next from CogSci 2025

Distinguishing Human vs. AI-Generated Texts: How Humor and Emotional Expression Shape Perceived Authorship
poster

Distinguishing Human vs. AI-Generated Texts: How Humor and Emotional Expression Shape Perceived Authorship

CogSci 2025

+1
Giulia Melis and 3 other authors

01 August 2025

Similar lecture

How Hypocritical Is Your LLM judge? Listener-Speaker Asymmetries in the Pragmatic Competence of Large Language Models
poster

How Hypocritical Is Your LLM judge? Listener-Speaker Asymmetries in the Pragmatic Competence of Large Language Models

ACL 2026

Judith Sieker and 1 other author

06 July 2026