JurisBench-Omni: A Judicial-Practice-Oriented Benchmark for Legal Omni-Modal Evaluation
Abstract
Large language models (LLMs) have achieved significant progress in *multi-modal* tasks; however, they still face challenges in the *omni-modal* understanding of real-world signals. Court hearings constitute one of the most consequential stages of judicial practice and are inherently omni-modal: judges, litigants, and witnesses jointly convey legal meaning through oral arguments, facial expressions, bodily actions, procedural contexts, spatial environments, causal relations, and other symbolic cues (e.g., silence). Much of this meaning is compressed into written judgments. Current legal AI evaluation remains predominantly text-centric or treats audio and video as separated and auxiliary signals appended to general-purpose multi-modal templates and, as a result, fails to assess whether models can comprehend judicial hearings in a human-like manner by integrating speech, actions, and intentions into a coherent, continuous task representation. We introduce JurisBench-Omni, a benchmark for legal omni-modal evaluation constructed directly from real hearing videos and their corresponding judgment documents. The benchmark decomposes judicial competence into four omni-modal dimensions: *role attribution*, *intention understanding*, *evidence recognition*, and *hearing-to-judgment grounding*.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.