Can LLMs Anticipate Their Cruxes?
Abstract
To use language models to help us make decisions we would like to understand the rationale behind the language model's outputs. In principle, the language model provides its rationale as part of its text output. However, the sheer volume of text makes this hard to use in practice. In this paper, we are interested in extracting the *cruxes* that lead a model to recommend one action over another. We take a crux to be a sub-question where changing its value may have altered the model's conclusion. We consider three strategies of increasing computational cost. First, we just ask a model to identify, a priori, what the cruxes of a decision will be. Second, we have the model actually go through the process of generating a response and then extract the cruxes that arose in its reasoning. Finally, we generate many responses and extract the cruxes by comparing the reasoning across the rollouts. We introduce a number of tricks to make this last procedure computationally tractable. We then conduct an extensive empirical evaluation to assess whether the computationally inexpensive options are sufficient. They are not. Across domains and models, we find that a model's a priori assessment of the cruxes has little relation to the factors that are actually identified as important a posteriori. Using a single rollout helps, but is still inadequate for identifying the full set of cruxes. Accordingly, for high-stakes decisions (and to understand the space of language model decision making) the many rollout procedure is unavoidable.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.