CogSci 2025

•

August 01, 2025

•

San Francisco, United States

keywords:

case studies

corpus studies

statistics

artificial intelligence

natural language processing

reasoning

In recent years, the remarkable performance of large language models (LLMs) in tasks such as legal judgment prediction (LJP) has garnered widespread attention. An increasing number of LLMs have been successfully implemented to assist judges in performing various legal tasks. However, their robustness and reliability in complex judicial scenarios remain a subject of debate, particularly when confronted with real-world legal cases. Existing research often overlooks the systematic evaluation of these LLMs in terms of judicial fairness, robustness and other ethical considerations. To fill this gap, we propose a novel benchmark that integrates authentic legal cases to evaluate the robustness of LLMs in the legal judgment prediction (LJP) task. Our work establishes foundational safety standards for applying LLMs in the legal domain.

Downloads

PaperTranscript English (automatic)

Next from CogSci 2025

Measuring the Semantic Consistency of Ordinal Annotations via Text Embedding Spaces and Its Applications
poster

Measuring the Semantic Consistency of Ordinal Annotations via Text Embedding Spaces and Its Applications

CogSci 2025

Yo Ehara

01 August 2025

Similar lecture

Multi-Defendant Legal Judgment Prediction via Hierarchical Reasoning
workshop paper

Multi-Defendant Legal Judgment Prediction via Hierarchical Reasoning

EMNLP 2023

Yougang Lyu
Zhumin Chen
+6
Yougang Lyu and 8 other authors

07 December 2023