acceptodds
Under review as a conference paper at ICLR 2027

HyperLogic: A Hard, Forward-Authored Chinese Logical Reasoning Benchmark with Execution-Derived Answers

Abstract

Existing logic benchmarks primarily measure models' ability to answer reasoning questions directly. Scalable benchmarks often generate text from formal structures, which makes answers easy to compute but fixes the formalization before the problem is written. Forward construction preserves the challenge of finding a faithful formalization, yet makes difficulty and answer reliability harder to control. We introduce HyperLogic, a forward-construction pipeline that separates problem authoring from answer generation. A multi-agent workflow hardens undergraduate-authored Chinese seeds without solving them; two agents from different model families independently translate each finished item into executable finite-domain models; their encodings and solver-derived answers undergo layered, agent-assisted adjudication under human-expert oversight. HyperLogic-Base contains items and sub-questions and separates seven frontier models by percentage points in strict item accuracy (–%). HyperLogic-Hard contains items with larger, coupled search spaces, on which no model exceeds % accuracy in direct answering. We also use Hard to evaluate agents' ability to formalize and solve problems with tools, comparing a code sandbox alone with one that includes our logic modeling library. The sandbox improves every model by – points; adding the library helps five models and hurts two. These results highlight the difficulty of faithful formalization even with tool access.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.