Language as Computation: Simulated Feedback and On-Policy Guidance for Interactive Agents
Abstract
We study the point at which language becomes not only a medium for supervising computation, but a medium for carrying out interaction. For interactive agents, language can serve not only as the environment interface but also as the medium of the learning interface. Because environmental signals and learning feedback are both expressed as text, they can be written directly into the interaction history and participate in subsequent action generation, without requiring access to teacher logits. We instantiate this idea as a language-mediated on-policy learning framework in which a fixed teacher produces action-conditioned environmental feedback, progress signals, and gradually fading hints for training a student policy with GRPO. We evaluate both the fidelity of the simulated feedback and the transfer of the resulting policy to executable environments. Across 5,868 replayed transitions, DeepSeek-V3.2 achieves 75.3% normalized observation exact match, compared with 67.9% for Qwen3-32B, while simulation fidelity decreases with interaction depth. A Qwen3-8B student trained through these interactions reaches 44.13% on BFCL Multi-Turn and 46.88% on Tau-Bench, exceeding matched executable-environment RL by 2.25 and 2.07 points, respectively; without hints, simulated-feedback training remains competitive on BFCL. The same framework transfers to GUI agents on WebArena, matching executable-environment training in three categories and improving CMS.These results suggest that language can unify environmental signals and learning feedback within the policy's computational context, where both directly shape subsequent actions.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.