ORIGATE: Diagnosing Provenance Invariance in Learned Action Reviewers
Abstract
An invariant pooling operator need not preserve the semantics of the objects placed inside it. ORIGATE diagnoses this distinction in action review by deriving paired tests from a declared provenance policy, inspecting intermediate representations, and intervening on a frozen encoding path. Neutral empty-slot handling removes sensitivity to repeating one source but does not equate a compound source with its partition across carriers. A local join intervention eliminates that partition discrepancy while preserving all trainable parameters, yet substantial policy errors remain. Matched-support fits further separate exposure to a source union's labels from exposure to its carrier encoding. When only split mixtures were trained, the intervention raises compound-input accuracy from about 80% to above 99.98%; when only compound mixtures were trained, it raises allowed-request false denial from near zero to 42.91% and 13.82% in two controls. A 192-request local study exercises these mechanism pairs through genuine context, process, and state entries while authorization remains independent. The results identify when encoding repair transfers a useful decision and when it transfers an erroneous readout. They concern trusted typed metadata and a finite policy, not language understanding or autonomous-agent safety.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.