Computational Understanding: Curriculum-Driven Construction and Provenance-Auditable Evaluation of Domain Reasoning
Abstract
Evaluations of LLM-based domain-knowledge systems are vulnerable to a quiet failure: the score can measure the substrate's parametric memory, or the evaluators' own leakage, rather than the constructed artifact. Auditing our own harness with a grounded rubric exposed two such channels — a hand-authored fact table and a construction cache — that supplied 13 of 50 development answers; we removed both, added specified controls, and rebuilt evaluation around a curriculum specification — a typed premise (syllabus), a corpus (textbook), and worked examples (problem sets) — scored by a deterministic certificate with the adaptivity of every later step disclosed under an explicit exposure taxonomy. On first evaluation of frozen, receipt-ruled, human-authored exams, the constructed stores lose to premise-conditioned retrieval-augmented prompting in both domains — 7/27 vs 17/27 (biochemistry) and 14/24 vs 15/24 (consumer-fraud law) — but every miss carries a provenance trail from which the builder diagnosed an addressable defect class. Across pre-registered repair iterations on those same fixed instruments (adaptive-reuse exposure disclosed throughout), the stores reach 20/27 and 23/24 with zero answer-time model calls and, under the stated confabulation definition, no invented text or sources identified in the audited campaign-final and fresh-exam outputs. We then ran a pre-registered zero-adaptivity generalization test: one fresh exam per domain, single scored runs. The fraud system met its pre-registered floor (10/13); the metabolism system failed decisively (2/19), and post-hoc diagnosis identifies a representation limitation — the fresh exam demands relation classes the practice-shaped record ontology never carried, while the engine's failure mode stayed typed refusals and grounded near-misses, with all designed probes refused. The curriculum specification, the provenance-auditable answer artifact, the disclosed repair loop, and the generalization test that binds its own authors are the contributions; the domain systems are the worked examples.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.