acceptodds
Under review as a conference paper at ICLR 2027

TI2V-Adapter: Efficient Adaptation from Text-to-video Models to Text-image-to-video Generation

Abstract

Text-image-to-video (TI2V) generation aims to synthesize realistic videos from text prompts and reference images. Existing deep-learning-based TI2V methods either require costly training with large-scale video data or rely on training-free inference-time manipulation of pretrained Text-to-video (T2V) models. While training-intensive methods require time-consuming optimization and substantial computational resources, training-free methods often exhibit limited performance because of the lack of learnable reference-image adaptation. These limitations motivate us to develop a TI2V method with low training cost and improved use of reference-image cues. We therefore propose TI2V-Adapter, an efficient adaptation method that introduces lightweight learnable modules to incorporate reference-image information into pretrained T2V models, enabling T2V-to-TI2V adaptation with minimal training cost. Specifically, TI2V-Adapter consists of a Visual Reference Adapter (VRA) that injects reference-image cues at both the input and feature levels of the pretrained T2V pipeline, and a Video Temporal Controller (VTC) that reinforces image information during denoising. Compared with training-free methods, TI2V-Adapter achieves a 15.3% FVD improvement with only 0.01% of the training samples and 0.83% of the trainable parameters required by training-intensive methods. Code is available in the supplementary material.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.