EEVEE: Post-Deployment Prompt Learning across Heterogeneous Tasks
Abstract
Prompt learning has become an important approach to adapting large language models to downstream tasks. However, some existing methods require access to model internals, while others are largely designed for single-dataset settings, limiting their practical applicability. In real-world deployments, models are often accessible only as black boxes, while incoming queries form heterogeneous streams drawn from multiple datasets, domains, and task distributions. We therefore propose EEVEE, a multi-dataset post-deployment prompt learning framework for LLM agents, enabling prompt learning under heterogeneous task streams. To mitigate cross-dataset interference, EEVEE introduces a router that partitions incoming inputs into task clusters and assigns them to suitable prompt configurations. This design is optimized via a router-prompt co-evolution strategy, which employs interleaved router and prompt learning phases to address their mutual dependency. Experiments across multiple datasets demonstrate that the framework improves robustness under heterogeneous data streams while maintaining single-benchmark learning capability and efficiency. Specifically, EEVEE improves average multi-benchmark scores by 10.38 and 24.32 points over Qwen3-4B-Instruct and DeepSeek-V3.2, surpassing SOTA methods GEPA and ACE by up to 37.2% and 48.2%.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.