Agentic Local Environment Alignment via Decision-Oriented Knowledge Learning
Abstract
Large language model (LLM) agents often operate in environments whose useful regularities are absent from task instructions: where objects are likely to be found, and which action procedures lead to success. These regularities can be learned from interaction, but trajectories also contain transient states and incidental actions that should not be carried across instances. We call regularities shared across related instances of a target setting local environment knowledge and organize them according to the decisions they support. Search decisions use probabilistic object-location relations, while execution decisions use procedural rules. We instantiate this representation in Structured Knowledge Learning (SKL), which keeps the LLM backbone frozen and learns both components from interaction: relations are estimated by combining commonsense priors with observation statistics across training instances, while rules are induced by contrasting trajectories with different returns. Both knowledge components are supplied to a frozen LLM agent at inference time. Under a shared evaluation protocol on ALFWorld and ScienceWorld, SKL substantially improves reward over prompting- and reflection-based baselines with GPT-4o-mini and Gemini-2.0-flash. On ALFWorld, for example, SKL improves reward from 62.6 to 81.3 with GPT-4o-mini and from 61.2 to 88.0 with Gemini-2.0-flash. Ablations show complementary gains from relations and rules, while budget and cost analyses show that these gains persist with fewer interactions and lower test-time token consumption. The learned knowledge also transfers across backbones: on ALFWorld, knowledge acquired with either GPT-4o-mini or Gemini-2.0-flash improves Gemini-2.5-flash, and GPT-4o-mini with SKL achieves higher reward than Gemini-2.5-flash without SKL.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.