URPO: A Unified Reward & Policy Optimization Framework for Large Language Models

Content not yet available

This lecture has no active video or poster.

AAAI 2026

•

January 24, 2026

•

Singapore, Singapore

Please log in to leave a comment

Downloads

SlidesPaper

Next from AAAI 2026

OmniEvent: Unified Event Representation Learning
poster

OmniEvent: Unified Event Representation Learning

AAAI 2026

+5
Zhipeng Cai and 7 other authors

24 January 2026

Similar lecture

Reward Model Evaluation via Automatically-Ranked Policy Alignment
technical paper

Reward Model Evaluation via Automatically-Ranked Policy Alignment

AAAI 2026

+1
Lei Ou and 3 other authors

22 January 2026