CETRA: Composable Evidence Certificates for Tool-Adaptive Risk Control of Multimodal Assertions
Abstract
Adaptive multimodal agents select tools using earlier observations, so calibration for one acquisition policy need not survive policy recomposition. We introduce CETRA, which converts tool outputs into action–history evidence increments for assertions fixed before acquisition. Cross-fitted calibration bounds logging-reference null means, while unsupported configurations contribute neutral increments. Under independent source-level calibration, stable tools, and explicit full-history conditional null-mean bounds, the product is a nonnegative supermartingale for each successful calibration artifact. Accounting for calibration failure yields anytime false-certification control across admissible acquisition and stopping policies. We evaluate CETRA in CETRA-Sim, ChartQA, and DocVQA. At , CETRA passes the policy-wise empirical confidence-bound checks on both real datasets, with worst-policy risk estimates and , while reusing one calibration object per dataset across seven policies. Coverage exceeds static certification by and percentage points. Policy-specific recalibration achieves higher coverage with separate full allocations, whereas CETRA achieves higher coverage under the evaluated matched total calibration-trajectory budgets. These results characterize the empirical coverage–reuse trade-off while keeping the assumptions required for theoretical certification explicit.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.