Don't Read the Log: Execution Traces Contaminate Verifiers in Video-Generation Agents
Abstract
Agentic video-generation systems close a loop between a generator and a verifier: an LLM plans shots, calls a text-to-video model, and a multimodal judge decides whether the result satisfies the request. To diagnose where a long workflow fails, recent harnesses deliberately show the judge more than the video—the agent's execution trace, its plan, the narration it synthesized. We ask whether this auxiliary text moves the judge's verdict on purely visual requirements, holding the frames fixed. On a benchmark of 109 generated two-event clips with manual labels, in which the requested event is either visibly completed or visibly missing (`near-miss' failures, the common case in practice), a trace that reports a successful tool call makes three open-weight Qwen-VL judges (7B, 8B, 32B) accept – of the failures, up from – without text, and a contradicting trace makes them reject up to of correct clips; an instruction to “use only the frames” does not remove the effect. Frontier closed judges (GPT-5.4-mini, GPT-5.5, Claude Opus 5, and DeepSeek-V4 reading a vision tool's description) are essentially unmoved on the same clips, showing that the vulnerability is a property of the judge's learned trust in tool logs rather than of the task. Plan-derived text carries no clip-specific information, so it can only shift a judge's operating point, and in a repair loop that shift becomes a cap on the true pass rate that no repair policy can exceed; the cap matches simulation to two decimals. In the loop, contamination is exploited without any adversarial agent: an honest LLM planner that always regenerates ends with a judge pass rate of and a human-labelled pass rate of , and a pipeline in which a cheap checker writes its verdict into the trace launders that checker's errors into a stronger final judge ( false accepts). We propose least-privilege judging: every requirement declares the evidence type that can satisfy it, and the judge sees only that evidence—frames for visual requirements, the trace for process requirements, with OCR-masking for text burned into the frames. It restores the true pass rate to – at the (necessary) cost of actually regenerating failed clips, loses nothing on process-level checks, and is a zero-cost guarantee for any judge, including ones whose trust in the trace is unknown.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.