acceptodds
Under review as a conference paper at ICLR 2027

Conferencing Tokens: Learning Compact Human Representations for Video Communication

Abstract

Video conferencing serves both human participants and intelligent assistants, yet pixel-based communication requires models to decode video and process long visual input sequences. We introduce Conferencing Tokens, a compact, transmissible representation of people in motion that connects the same decoded output to different downstream models. Initialized from pretrained geometry, the tokenizer combines rate–distortion learning with causal temporal entropy modeling to form a shared framewise interface. The decoded representation supports video reconstruction with person-specific appearance models and enters a language model directly through a learned interface, using one input position per frame. Experiments show a mean BD-PSNR gain of dB over a facial-parameter-based conferencing coding baseline. For transcript-conditioned meeting-state prediction, language adaptation retains accuracy close to an unadapted dense RGB reference while reducing modality input positions by a factor of . These results demonstrate the feasibility of reusing a compact communication representation directly for downstream analysis, linking efficient transmission with model consumption.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.