One Candidate Changes the Rest: Non-Local Instability in LLM-based Ranking
Abstract
Large language models (LLMs) are increasingly used to recommend and rank candidates in a variety of settings, including e-commerce and hiring. In these LLM-as-a-Recommender scenarios, models are often tasked with mapping an unordered set of candidates to an ordered ranking. While prior work has shown that LLM-based rankers can be sensitive to presentation order, we identify a distinct and previously underexplored property of these rankers: Non-Local Instability. Ideally, a non-locally stable ranker should ensure that modifying a single candidate affects only that candidate's position relative to the others. However, we show that current autoregressive LLMs systematically violate this property: small, localized changes to the description of a single candidate can alter the relative ordering of other candidates whose descriptions remain entirely unchanged. We characterize this phenomenon across multiple types of perturbations, evaluate its prevalence across different models, and investigate its underlying mechanisms. Our experiments show that non-local instability is widespread and persists even in recent state-of-the-art models. Reasoning can substantially mitigate this effect, but at the cost of increased ranking variability under the same context. These findings reveal a fundamental robustness and fairness concern in LLM-based ranking systems: the relative ranking assigned to a candidate can be substantially influenced by seemingly irrelevant semantic changes to other candidates.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.