MosaicDrive: Task-Aware Multi-Camera Reasoning for Autonomous Driving
Abstract
Tool-augmented driving agents typically acquire additional camera views when they are uncertain about what they have already seen. We identify a structural limitation of this design: the confidence that triggers retrieval is computed from evidence the agent already holds, and so carries no information about regions it has never observed. Acquiring more evidence also introduces a second difficulty: a newly retrieved observation can contradict or override a correct prior judgment, so broader acquisition is not monotonically beneficial. We therefore treat acquisition and verification as parts of a single decision. We present MosaicDrive, a driving agent framework that derives camera and tool requirements from the question, requires the selected evidence before producing an answer, and resolves the resulting observations through Arbiter, a source-aware verifier that checks claims against available evidence, issues targeted re-queries, and preserves conflicts that cannot be settled. We evaluate on MOSICQA, a multi-camera driving question-answering resource with executed tool traces and scene-disjoint splits, under a protocol that separates coverage from selectivity, correction from abstention, and answer quality from acquisition cost. Controlled camera-policy comparisons and evidence-corruption tests isolate the contribution of task-conditioned acquisition and of explicit verification.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.