acceptodds
Under review as a conference paper at ICLR 2027

CertiHop: Certified Necessary-Evidence Supervision for Long-Context Reasoning

Abstract

Long-context question answering benchmarks often annotate supporting evidence, yet support is weaker than necessity: a model may answer correctly without combining every designated item. This gap grows when a compact proof is expanded into long documents, where redundancy, alternative paths, answer leakage, and parametric knowledge can reduce effective dependency order. We introduce CertiHop, a construction and certification framework for long-context necessary-evidence reasoning. Each example begins with an executable dependency tree, expands into a multi-document dossier, and is certified only on the final context. Full must uniquely determine the answer, whereas Zero and every required-leaf leave-one-out intervention must preserve the gold answer as possible but not uniquely determine it. The resulting family supplies paired sufficient/insufficient supervision and parent-level Joint Irreducibility (JI); non-critical interventions distinguish selective sensitivity from generic abstention. Across four Qwen-series backbones, CertiHop yields LongBench-3 point estimates 11.92–36.25 points above equal-size continued-normal training. The controlled A3 comparison yields positive LongBench-3 differences at all four scales. A 9B RULER decline exposes a checklist-completeness boundary. CertiHop turns intended graph dependency into an auditable property of the final text.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.