acceptodds
Under review as a conference paper at ICLR 2027

Aligning to New World Models Using Constitutional AI

Abstract

As large language models (LLMs) have scaled, so has the quality of their implicit world models: an internal understanding of an environment. However, there are environments in which even frontier LLMs don't work correctly out of the box, which we call new world models. We ask how to align a frozen LLM with a new world model using a small set of solved training problems and no weight updates. We build five such environments from Manhattan routing and orbital mechanics, on which a frontier LLM fails to get more than half of the evaluation correct. We use sleep-time compute: before test time, the model solves training problems, checks its answers against a simulator, and revises a constitution, a list of single-clause rules that we place in its prompt. While prior work uses constitutions to align models with human values, we repurpose them to instead state the governing rules of an environment. We add a fixed slot budget to makes rules compete for space, and an ensemble procedure to account for variance across independent runs. Ensembled constitutions improve accuracy by 1.8 over zero-shot and outperform few-shot prompting, an RL-trained advisor, and Inverse Constitutional AI. An ensembled constitution also slightly outperforms unconstrained prompt optimization (GEPA) with a 3.9 shorter prompt. Taken together with the advantages of constitutions, such as interpretability and editability, our results show that constitutions are an effective method of aligning frozen LLMs for new world models.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.