AgenticVBench-Omni: Evaluating Omnimodal Agents on Economically Valuable Work
Abstract
We introduce AgenticVBench-Omni, a benchmark that tests whether omnimodal agents can turn video, audio, images, and documents into economically valuable work. Its 45 tasks span eight families and work activities linked to 10 occupations, requiring precise records, quantitative annotations, editable assets, or completed environment states. Designed around complex professional workflows, the tasks require sustained evidence gathering, cross-modal reasoning, and tool use while preserving temporal, spatial, and identity relationships. Each task specifies an output contract and uses a fully deterministic verifier to compute scores, without rubric-based grading by a model. Across five frontier agents evaluated three times per task, the strongest agent achieves a mean score of just 0.283, and 60% of attempts score below 0.10. Verifier outputs and execution traces show that agents recover partial information but struggle with precise localization, complete coverage, accurate reconstruction, and self-verification of the final deliverable.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.