COKE: A Unified Image-Video Encoder Trained on Public Corpus
Abstract
Leading vision encoders often rely on training corpora that are not publicly released, limiting access to comparable training resources. We present **COKE**, a family of unified image-video encoders and a training recipe using public data and priors from released models. The family includes COKE-ViT-H (0.6B) and COKE-ViT-G (1.9B), with shared image-video weights at each resolution. ***Visual Reference Pretrain*** applies auxiliary visual supervision through dedicated reference-token outputs and improves early sample efficiency. At a fixed distillation weight, this pathway achieves higher classification and retrieval scores than pooled-patch distillation. ***Snapshot-Replay*** periodically interpolates video-trained weights with a fixed image-pretrained checkpoint and resumes optimization from the fused weights. In controlled COKE-ViT-H-256 experiments, it improves image and video averages over continuous video-only fine-tuning by 1.81 and 1.03 points, respectively, without replaying image-text data. A resolution curriculum supports fixed and native inputs, while COKE-K and COKE-V guide concept balancing alongside quality filtering and deduplication to construct the 1.68B-pair **COKE-1.7B** corpus. Fixed-budget ablations support COKE-K vocabulary selection and the utility of a data mixture incorporating COKE-V-selected samples. At 512 resolution, COKE-CLIP-G achieves 86.8 average zero-shot accuracy across six image classification benchmarks and 79.9 average retrieval Recall@1, compared with 86.6 and 78.9 for PE_core-G under our evaluation protocol. Frozen-encoder probing and multimodal evaluations further demonstrate the utility of the learned representations beyond zero-shot image-text tasks. We will release our models, vocabularies, and evaluation code.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.