acceptodds
Under review as a conference paper at ICLR 2027

CapSieve: Coupled Modality, Extent, and Stopping under Multimedia Serving Budgets

Abstract

A multimedia request can begin with a cheap transcript, expand into dense frames, approach its decode and context limits, and still cross the wall-clock deadline despite every acquisition being plan-feasible when selected; stopping earlier can instead force abstention or an unsupported answer. CapSieve controls these competing exits with one masked policy over modality–extent–stop tuples and a fail-closed answerability gate under latency, decoded-media, and token budgets. An executed-route ledger links every answer or abstention to native task quality, realized SLA compliance, acquisition work, and claim support. Across DocVQA, ChartQA, ActivityNet-QA, and AudioCaps, CapSieve scores 71.20.8 at the 4 s tier versus 72.30.9 for uncapped fetch-all, while reducing median normalized decode work from 1.00 to 0.39 and passing the SLA on 95.2% of the fixed mixture. It exceeds a matched constrained actor-critic by 0.6 points (95% CI [0.2, 1.0]), improves automatic evidence alignment by 3.5 points, and uses 0.04 less median decode; reward-only training reaches 70.8. A blinded 4 s audit independently yields 86.2% citation-support precision and 79.1% recall, compared with 80.3% and 76.4% for the actor-critic. Route accounting then locates the shared serving pressure: visual-temporal and multi-hop requests occupy 33.1% of the mixture but carry 59.7% of decode work, 66.7% of abstentions, 74.5% of SLA failures, and 74.2% of evidence contradictions—up to 2.25\(\times\) their prevalence. The pattern persists across cap tiers, policies, and a second VLM backbone, turning executed routes into a concrete guide for where additional serving and modeling capacity has the highest leverage.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.