acceptodds
Under review as a conference paper at ICLR 2027

Behavior Gradient Fields

Abstract

For a given model behavior, e.g., out-of-domain generalization or refusal of harmful requests, many models may exhibit this desired behavior, and there are thus many datasets which could have produced these models. On one hand, the multiplicity of permissible datasets is useful; understanding this space of datasets and their induced models can inform how robust a target behavior is, where it conflicts with other objectives, and whether alternative paths avoid those conflicts. However, the known subset of *inducing datasets* is typically small, with a few datasets either given or produced through simple data-mixing procedures. To this end, we propose **behavior gradient fields**, a framework for tracing model behavior and its geometric implications across data, weight, and behavior spaces; we introduce practical algorithms for exploring these fields via **data steering maps**. Within this framework, we investigate spurious correlations in image classification, as well as steering refusal and sycophancy in LLMs. We use interactions between behavior fields to understand when objectives align or conflict, explain unexpected side effects, and jointly steer behaviors. We also find that some fields are more concentrated, with implications for robustness, and that we can find paths around apparent disjointness in behavior fields. Our results emphasize the consequential properties of the geometry of inducing datasets.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.