Fragility is not Reducible to Hardness: Conditional Causal Difficulty for LLMs
Abstract
Large language models show large, seemingly arbitrary performance drops when a task is rewritten in ways that leave its requirements untouched. We recast task difficulty as a causal quantity: the treatment effect of meaning-preserving perturbations on a model. A meaning-preserving rewrite can only be observed to hurt a task the model can already solve, so we place the estimand on the principal stratum where the degradation occurs: the conditional causal difficulty (CCD), the probability that a randomly drawn twin breaks a solvable task. Because harder tasks tend to be more phrasing-sensitive, complexity enters as an effect modifier rather than a confounder, and we residualize it out to obtain complexity-residualized difficulty (CRD); zero for a task exactly as fragile as its complexity predicts, positive for anomalously fragile ones. Current benchmarks do not supply the data CCD requires, so we introduce GSM-IC-Adv and MATH-Adv, augmenting GSM-IC and MATH with twins built from perturbation families, raising fragility on GSM-IC-Adv by over the original GSM-IC () while keeping the answer key unchanged. Our predictor learns the per-perturbation break probability and reaches of the suite reliability ceiling on GSM-IC-Adv and on MATH-Adv, against / for a surface-feature baseline and / for supervised baselines such as VAP3, and doubles random recall of fragile items under a curation budget. Additionally, cross-fitted residualization leaves CRD uncorrelated with the declared proxy in practice.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.