Planted Effects Give Detection Floors for Reasoning-Trace Measurements
Abstract
Studies that intervene on a model's reasoning trace report how strongly its answer depends on the trace, but rarely the smallest effect their design could detect. Without that detection floor, a reader cannot tell a small reported change from item-sampling noise. We measure the floor for the causal importance of reasoning (CIR) by planting effects of known size in an open model's own reasoning. The answer must repeat a code word shown at the end of the reasoning, and switching that word on a chosen share of items plants the effect. The spread of the planted effects across items gives the floor. It shrinks with the square root of the item count, which also tells a new study how many items a planned effect needs. In a published reinforcement-learning study, and assuming item variance carries over from our calibration model, 34 of 36 task effects are resolvable. The exceptions are the study's smallest effects, including a near-zero change its design could not have resolved. Studies that report CIR effects should report this floor with them, and planting effects of known size in the studied model gives it.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.