Where Is the Seahorse Emoji? When Language Models Fail to Settle
Abstract
Ask a language model for the seahorse emoji, and it may enter a prolonged cycle of proposing, rejecting, and reconsidering nearby symbols instead of correctly saying that no answer exists. We call this behavior failure-to-settle. We find that it extends beyond emojis to rare-language vocabulary and scientific and standardized symbols, suggesting a broader failure mode in reasoning and retrieval over externally defined symbolic systems. Across multiple open-weight model families, we show that failure-to-settle can be causally controlled through interventions on model internals and output representations. Steering along a direction associated with belief in an item’s existence substantially reduces failure-to-settle. Similarly, steering toward a reasoning-persistence direction can reduce failure-to-settle by promoting commitment to a final answer, whereas steering toward continued reasoning increases it. Finally, making a valid output unavailable can induce failure-to-settle, while constructing a synthetic output token can prevent it. These results distinguish failure-to-settle from generic repetition and show that interpretable interventions can directly control whether a model continues searching or settles on an answer. We introduce a benchmark for measuring this behavior and find that it remains common in state-of-the-art open-weight models.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.