acceptodds
Under review as a conference paper at ICLR 2027

TAMER: Testing Agents in complex Multi-tool Enterprise knowledge worker Roles

Abstract

We introduce TAMER, a benchmark for evaluating vision-language-model agents on realistic knowledge work in an enterprise setting. TAMER contains 238 tasks requiring an agent to coordinate actions in a simulated environment consisting of eight web apps it observes only as screenshots and operates through coordinate- based tools, grounded in a synthetic organization of 54 employees across 19 teams and 21 projects spanning five months of interlinked history. Each task ships a hybrid verifier combining deterministic scripts with LLM-judged rubrics. Measured difficulty splits the set into 24 easy, 98 medium, and 116 hard tasks. Across six models, the highest per-attempt accuracy on the full set is 42.3 ± 3.6%, falling to 6.0% on the hard subset. Pass rate also falls with the number of web apps a task spans.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.