acceptodds
Under review as a conference paper at ICLR 2027

CoWA: Learning Compact Decision Policies from Controller-Executed Workflows

Abstract

Large language models (LLMs) are moving from open-ended reasoning toward constrained system control, where agents transform structured states into controller-executable actions that optimize system utility under operational constraints. This shift calls for compact LLM agents that can operate under tight latency, cost, and privacy requirements without repeatedly querying cloud models. Existing methods study tool interaction, trajectory supervision, and post-training adaptation, but leave a key design question unresolved: which parts of a domain workflow should be executed by the controller, and which should be learned by the model? We propose Cold-start Workflow Alignment (CoWA), a framework that separates fixed domain computation from final-action learning. CoWA performs task and interface alignment, tool-augmented cold-start supervision over a fixed controller-executed workflow, and GRPO-based objective alignment. The compact policy learns to synthesize the final action from tool-derived observations, while the controller retains responsibility for tool execution and post-generation checks. This design supports constraint-aware decision-making without iterative per-instance utility search. We instantiate CoWA in a representative Wi-Fi channel-allocation case study. Using 750 training trajectories, a CoWA-trained Llama-3.1-8B achieves an average throughput of 293.5 Mbps, compared with 291.8 Mbps for a prompt-engineered GPT-5.2 baseline, with an observed end-to-end wall-clock speedup in our evaluation setup. CoWA also approaches the performance of the search-based baseline without using its search budget.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.