TraceSAE: Selective Sparse Control After Tool Feedback
Abstract
Tool-using language models must revise plans when tool feedback invalidates their prerequisites, while retaining plans that remain supported. We introduce TraceSAE, an inference-time controller that targets unwarranted persistence: continuing to rely on a premise explicitly contradicted by visible tool feedback. TraceSAE separates the decision to intervene from the action it changes. An unexecuted shadow proposal informs a gate that is frozen before the actual action is generated. When the gate opens, the controller edits 16 sparse-autoencoder coordinates toward activation profiles from warranted decisions, preserves the original reconstruction error, and caps each perturbation at 2% of the residual norm. Feature groups and intervention strengths are selected through restored-state continuations that jointly evaluate action correctness, goal progress, preservation, and token cost. The language model and pretrained autoencoder remain fixed. Across Gemma-2-9B-it and Llama-3.1-8B-Instruct on Telecom-Compose and ALFWorld, TraceSAE improves full-task success over a rank-16 dense controller by 4.67–8.21 percentage points under a shared gate, donor bank, edit window, and candidate-search allowance. These results establish the task-level benefit of behaviorally selected sparse-profile control in the four evaluated settings.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.