LIUYAO-Bench: A Hierarchical Benchmark for Expert-Domain Verification in Classical Chinese Texts
Abstract
Liuyao is a rule-governed expert-text tradition in classical Chinese where judgments depend on hexagram configurations, calendrical context, line relations, and transformations. Evaluating large language models (LLMs) in this domain requires examining both their knowledge of individual rules and their ability to verify interpretations of specific cases. We introduce LIUYAO-BENCH, a hierarchical benchmark containing 3,060 binary verification instances drawn from four classical Liuyao texts. The benchmark organizes evaluation into three layers: (1) basic knowledge, covering concepts, relations, and rule statements; (2) intermediate-state interpretation, covering yongshen selection, line strength, and dynamic line relations; and (3) full-case judgment verification, checking candidate final judgments against structured case records. All three layers ask models to assess a supplied candidate, allowing comparison across different scopes of verification. Across ten contemporary LLMs, the best model reaches 79.77% overall accuracy, while dynamic relations and complete cases remain difficult. On a balanced 300-item subset, two practitioners achieve 88.67% mean accuracy, compared with 79.00% for the best model on the same items. Both practitioners outperform all evaluated models on Dynamics and Full-Case. Additional experiments with Qwen3-8B show that chain-of-thought prompting improves overall benchmark accuracy, while fine-tuning improves overall accuracy on a held-out source book; full-case performance remains limited in both settings. LIUYAO-BENCH provides a structured setting for examining the gap between local rule knowledge and case-level verification in a classical Chinese expert domain.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.