acceptodds
Under review as a conference paper at ICLR 2027

Live Assistant: Learning Whether, When, and Whom to Assist in Real-World Live Social Streams

Abstract

Livestreams are a primary form of online video: long-lasting, interactive environments in which omni-modal content, viewers, hosts, and platform signals unfold together. Yet how AI systems should act within them remains largely unexplored. The central question is how a model can move beyond perceiving what is happening and deciding when to respond proactively, toward selective participation in a shared social environment. We introduce \liveassistant, a framework for mixed-initiative, role-conditioned assistance that formulates livestream assistance as four coupled decisions over a multi-party stream: whether to act, when to act, whom to address, and what to communicate. At each incoming chunk, \liveassistant integrates native audio and video with synchronized comments, gifts, viewer dynamics, and room metadata, and chooses to continue observing, update private memory, or produce a grounded response for viewers or hosts. To learn this policy, we build a streaming trajectory engine that reconstructs replay timelines into time-aligned decision trajectories, yielding over 320 hours of training data and a 39-hour human-verified benchmark spanning 8 regions and 7 languages. We further introduce Marker-Aware Multiturn Supervised Fine-Tuning (MA-MSFT) and Streaming Multiturn GSPO (SM-GSPO) to strengthen structured decisions and improve behavior over evolving streams. \liveassistant attains the strongest joint performance in timing, routing, and content quality among evaluated models, moving streaming assistance from reactive responses toward selective participation in live social environments.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.