VLMintune: A Unified Interface for Developing Visual Instruction Tuning Methods
Abstract
Developing a visual instruction-tuning method often requires changes beyond model parameters, including inserted modules, forward hooks, token masks, supervision, objectives, and inference reconstruction. When these changes are scattered across paper-specific code, reusing a method requires substantial integration effort. We present VLMintune, a research library that treats the complete tuning method as the reusable unit. Each tuning recipe specifies how a method modifies a supported VLM, what it trains, any method-specific inputs, supervision, or loss, and how its learned behavior is saved and reconstructed. The surrounding data processing, backbone-native formatting, training, inference, and evaluation remain shared. VLMintune supports weight adaptation, inserted modules, representation interventions, and alternative token supervision. We implement seven representative methods and evaluate them on Qwen2.5-VL-3B-Instruct across TextVQA, VizWiz, and ScienceQA, yielding 21 full-data runs that demonstrate consistent within-platform comparison of structurally different tuning methods.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.