Invarialoop: Operator-Relative Verification for Multi-Route Inference
Abstract
Multi-route inference systems often map, screen, and combine several candidates before returning one output. Verification can then be mis-scoped: a check on the candidates or an intermediate object is treated as evidence about a different downstream output. We call this operator-relative verification and evaluate it by holding candidate realizations fixed while changing only the downstream operation or verifier. Controlled studies show that a mean constraint can hide large violating mass and that heterogeneous structured outputs require canonical mapping before cross-route checks are meaningful. On a public hourly Bike Sharing dataset with 17,379 records, five regressors predict casual rentals, registered rentals, and total rentals under the exact identity cnt=casual+registered. Screening reduces the empirical mass that violates the identity from 19.8% to 0.7% in a random split and from 26.0% to 0.5% under a 2011-to-2012 shift; after screening, the mean and violation-mass verifiers make the same decisions. On 100 held-out balance-sheet images from United States Securities and Exchange Commission (SEC) Form 10-Q filings, three vision-language model routes extract Assets and Liabilities-and-Equity while eXtensible Business Reporting Language (XBRL) facts are withheld until scoring. A screened arithmetic average satisfies the accounting identity on all 99 filings it releases but matches both withheld totals on only 56; the intact Qwen3.5-9B constituent matches 94 on those same filings. Among the paired cases, 38 are correct only under Qwen3.5-9B and none are correct only under the average. In a separate study with eight samples per route, with the candidates and arithmetic aggregation fixed, a violation-mass verifier releases one output whose arithmetic mean violates the accounting identity by $103,878.94, while a mean-residual verifier abstains. These results motivate reporting the candidate set, downstream operator, verified property, final output, and held-out task correctness separately.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.