CogSci 2025

•

August 02, 2025

•

San Francisco, United States

keywords:

computational modeling

decision making

learning

psychology

Life often presents choices that are not mutually exclusive, yet there has been insufficient research on human learning and directed exploration involved in combinatorial settings. We investigated human behavior in a four-armed combinatorial bandit (CB) task (N=107) where participants combined "nutrients" affecting required nurture time of virtual plants. Participants demonstrated effective learning but converged to suboptimal strategies, preferring combinations of one or two options. To model learning, two computational models were proposed and compared: a naïve extension of upper confidence bound (NaiveUCB), and a linear UCB model (LinUCB), both incorporating heuristic components. The NaiveUCB model with penalty for multiple selections, value decay, stickiness, and recency-based credit assignment best explained behavior, outperforming both LinUCB and simplified variants, suggesting that humans may navigate uncertainty through simple heuristics rather than sophisticated estimation. These findings extend our understanding of exploration and credit assignment in CB, and provide insight into daily decision making.

Downloads

Paper

Next from CogSci 2025

VGG-19 Displays Human-like Biases in Statistical Judgment from Visual Graphs
poster

VGG-19 Displays Human-like Biases in Statistical Judgment from Visual Graphs

CogSci 2025

+1
Ruiyi Ding and 3 other authors

02 August 2025

Similar lecture

Revisiting the Role of Uncertainty-Driven Exploration in a (Perceived) Non-Stationary World
poster

Revisiting the Role of Uncertainty-Driven Exploration in a (Perceived) Non-Stationary World

CogSci 2021

Dalin Guo
Dalin Guo and 1 other author

27 July 2021