EthoBehavior: Learning Compositional Ethogram for Primate Behavior Understanding
Abstract
Primate behavior understanding seeks to identify and temporally characterize behavioral events and states from video, providing scalable measurements for research in cognition and sociality. Existing methods operate within independent ethograms whose categories differ in granularity and annotation scope, limiting the comparability of behavioral measurements across studies. Although Video-Language Models (VLMs) capture fine-grained visual cues, they typically treat ethogram categories as atomic concepts and remain agnostic to the structured behavioral semantics. To address these limitations, we introduce Compositional Ethogram Factorization, a shared representation that models primate behaviors as structured combinations of reusable behavioral predicates and typed attributes. We further instantiate it in PrInstruct-34K, a multi-task video instruction dataset comprising 34K examples from four primate behavior benchmarks. It retains 58 ethogram categories and standardizes their operational definitions into three complementary tasks: multiple-choice question answering, video captioning, and multi-label temporal parsing. Building on this, we develop EthoBehavior, a VLM with an Ethogram Composition Adapter that estimates factor distributions from video under partial supervision and encodes them as continuous Ethogram Factor Tokens for free-form behavior description and structured behavior analysis. Experiments across all three tasks and on unseen factor compositions show that EthoBehavior consistently outperforms instruction-tuned baselines and large open-source VLMs including GLM-5.3-Flash and DeepSeek-V4.1-Flash. Together, these results establish compositional ethogram representations as a principled basis for generalizable and comparable primate behavior analysis.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.