acceptodds
Under review as a conference paper at ICLR 2027

TaxLogic: Executable Rule Constraints for Rule-Intensive Professional Reasoning

Abstract

Answers in taxation and auditing are governed by explicit written rules, namely statutory rates, arithmetic identities, mandatory evidence, and prescribed procedures . Two bottlenecks arise when language models are deployed in such settings . First, general retrieval augmentation draws on heterogeneous open corpora, so conflicting or outdated provisions are intro duced and the uniqueness and authority of the knowledge source required by rigid business processes cannot be guaranteed . Second, conventional knowledge fusion encodes rules implicitly in model parameters, so no hard constraint exists along the inference path and intermediate steps violating the underlying business logic survive into an output that appears plausi ble while failing the compliance criterion . A two-layer constrained gener ation paradigm is proposed in which a rule engine is placed upstream of generation and an authoritative textbook corpus serves as the knowledge base . Compliance logic, arithmetic relationships, and required elements are compiled into executable predicates, and each option of a candidate answer must cite either a rule identifier or a textbook page together with a verbatim quotation, with every quotation located against the complete corpus rather than a retrieved fragment . A candidate in which any option fails is voided, and an item that neither source can adjudicate is referred for human review . Closed-form violation bounds are derived for four strategies . Certification and retry-based post-hoc checking are shown to share one lower bound, namely the share of errors the encoding cannot see, and regeneration can not reach it because a retry carrying a blind-spot violation is accepted and delivered . Two benchmarks with disjoint roles are evaluated . On 231 past CPA examination items, comprising 126 taxation and 105 auditing items with gold labels from the official answer keys, an open 27.9B-parameter model violates the governing rule on 43.7% of items without material, on 18.6% with textbook material only, and on 14.3% with both sources under the citation contract . The reduction over the textbook-only baseline is 4.3 percentage points and is not significant (McNemar p = 0 .14), a narrowing attributed to the compilation ceiling rather than mechanism failure . Every delivered answer carries a verbatim locatable quotation, 217 of 217 or 100%, while the entry-level location rate is 99 .1%, or 874 of 882 . On a 360-item rule-derived benchmark that invokes no language model and injects known violations into a shared perturbation stream, TaxLogic retains 6 .1% vio lations against 21.7% for retrieval augmentation and 11.7% for single-pass post-hoc checking, whereas ten retries still retain 8 .1%. The architecture is decoupled from the model backbone, requires no parameter fine-tuning, and is offered with a reproducible benchmark framework and an explicitly bounded set of trustworthiness metrics

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.