OmniCens: Self-Improving Distillation for Multimodal Classification of Millions of Profiles
Abstract
Classifying millions of social-media profiles from images, text, and contextual evidence is valuable for audience measurement but expensive when every prediction depends on a frontier model. We propose OmniCens, a self-improving distillation framework for reducing this dependence through local open-source multimodal models. Its central principle is that an expensive teacher response can serve both the current classification and subsequent student learning. We formulate the objective as minimizing amortized classification cost under precision, recall, and coverage constraints, accounting for teacher inference, calibration, student updates, and local execution. A confidence-based acquisition rule directs teacher supervision toward uncertain profiles, and repeated distillation targets recurring weaknesses while retaining the complete input evidence and attribute schema. The evaluation protocol isolates the value of continued learning and selective supervision through comparisons with fixed distillation, non-learning cascades, and random acquisition under matched budgets. Independent annotations assess predictive quality separately from teacher agreement, while held-out profiles and communities assess generalization. The key question is when reductions in teacher dependence recover the cost of adaptation without degrading classification quality. OmniCens connects this trade-off to a practical objective: making detailed audience characterization and creator–brand matching feasible for researchers and organizations with limited computational resources.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.