acceptodds
Under review as a conference paper at ICLR 2027

Omni-CoWork: From Collaborative Audio-Visual Discussions to Workspace Edits

Abstract

Large language model (LLM) agents can substantially support collaborative work by assisting with coding, document editing, and content creation. However, when teams discuss changes around shared artifacts, users must still consolidate their evolving decisions into explicit instructions before delegating edits to agents, adding effort and delaying execution. This raises a key question: can agents infer and implement a group's final decisions directly from audio-visual discussions into workspace edits? To explore it, we introduce Omni-CoWork, a benchmark for translating collaborative audio-visual discussions into verifiable workspace edits. It contains 124 tasks with human-recorded views and 1,213 synthetic tasks, with executable checks assessing both required changes and preservation constraints. Across seven agents, task success reaches at most 4.0% and 24.3% on the two splits, respectively, whereas substantial gains with gold instructions point to intent understanding as a major bottleneck. Investigating this bottleneck, we find that richer audio-visual descriptions alone remain insufficient: task success improves when agents ground discussion references in artifacts and revisit evidence during execution. These findings highlight the need for grounded collective reasoning: recovering final collective decisions, grounding them in artifacts, and verifying their realization as the workspace evolves. We translate these requirements into an Omni-CoWork skill that yields consistent gains on two agents but leaves most tasks unsolved, underscoring the need for more effective methods to bridge the gap to reliable practical use.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.