SoftSkill: Behavioral Compression for Contextual Adaptation
Abstract
Natural-language skills let agents reuse task knowledge, but long skill documents must be reinterpreted on every call. We introduce SOFTSKILL, which converts a skill document into a compact trainable context by initializing virtual-token embeddings from its text and optimizing an additive update with next-token prediction while freezing the language model. On Qwen3.5-4B, a 32-token soft skill improves over no skill by 7.6, 42.1, and 1.3 points on SearchQA, LiveMath, and DocVQA, respectively, and exceeds the optimized textual skill on SearchQA and LiveMath by 4.5 and 12.5 points while replacing artifacts containing hundreds to thousands of tokens. The same interface also improves multi-step execution: on Qwen3.6-35B-A3B, OfficeQA and ALFWorld improve over no skill, while in a controlled Qwen3.5-4B ALFWorld setting, training raises success from 44/134 to 91/134. Mechanistic interventions show that this gain is carried by a small, task-aligned continuous displacement: moving halfway along the learned direction recovers most of the improvement, an equal-norm random displacement does not, vocabulary projection removes the gain, and early-layer state replacement strongly suppresses it. These results show that task supervision can compile a readable skill artifact into a compact continuous control for both answers and actions without updating the backbone.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.