acceptodds
Under review as a conference paper at ICLR 2027

LIUYAO-Bench: A Hierarchical Benchmark for Expert-Domain Verification in Classical Chinese Texts

Abstract

Liuyao is a rule-governed expert-text tradition in classical Chinese where judgments depend on hexagram configurations, calendrical context, line relations, and transformations. Evaluating large language models (LLMs) in this domain requires examining both their knowledge of individual rules and their ability to verify interpretations of specific cases. We introduce LIUYAO-BENCH, a hierarchical benchmark containing 3,060 binary verification instances drawn from four classical Liuyao texts. The benchmark organizes evaluation into three layers: (1) basic knowledge, covering concepts, relations, and rule statements; (2) intermediate-state interpretation, covering yongshen selection, line strength, and dynamic line relations; and (3) full-case judgment verification, checking candidate final judgments against structured case records. All three layers ask models to assess a supplied candidate, allowing comparison across different scopes of verification. Across ten contemporary LLMs, the best model reaches 79.77% overall accuracy, while dynamic relations and complete cases remain difficult. On a balanced 300-item subset, two practitioners achieve 88.67% mean accuracy, compared with 79.00% for the best model on the same items. Both practitioners outperform all evaluated models on Dynamics and Full-Case. Additional experiments with Qwen3-8B show that chain-of-thought prompting improves overall benchmark accuracy, while fine-tuning improves overall accuracy on a held-out source book; full-case performance remains limited in both settings. LIUYAO-BENCH provides a structured setting for examining the gap between local rule knowledge and case-level verification in a classical Chinese expert domain.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.