Distractor-Based Jailbreaking Attacks in Language Models and Associated Changes in Chain-of-Thought Content (Student Abstract)

Content not yet available

This lecture has no active video or poster.

AAAI 2026

•

January 23, 2026

•

Singapore, Singapore

Please log in to leave a comment

Next from AAAI 2026

How Reasoning Influences Intersectional Biases in Vision–Language Models (Student Abstract)
poster

How Reasoning Influences Intersectional Biases in Vision–Language Models (Student Abstract)

AAAI 2026

Mohna Chakraborty and 2 other authors

23 January 2026

Similar lecture

Robust Safety Classifier Against Jailbreaking Attacks: Adversarial Prompt Shield
workshop paper

Robust Safety Classifier Against Jailbreaking Attacks: Adversarial Prompt Shield

NAACL 2024

Jinhwa Kim
Ali Derakhshan
Jinhwa Kim and 2 other authors

20 June 2024