SCALE: Controlling Compositional Capability Gains in Black-Box LLMs
Abstract
An assistant can withhold a complete answer while still releasing the pieces needed to reconstruct it. We study capability control over adaptive interactions, where the goal is to preserve permitted assistance while limiting an attacker's final task success. We propose , a controller that learns how much a candidate response changes task solvability given the current interaction history. Paired reconstruction attempts supervise before- and after-release solvability predictions, and an explicit difference loss emphasizes the candidate's contribution. Training pairs that hold the candidate fixed while varying earlier releases teach the auditor how this contribution depends on history. At deployment, screens candidates by predicted cumulative solvability and selects useful help with a penalty on positive predicted gain. Across Qwen3-8B and Llama-3.1-8B-Instruct on mathematics and code, retains 67.2–71.9% permitted-task utility at 20.6–21.8% attack success. Its utility exceeds that of the history-aware THRD baseline by 12.8 percentage points on average, while attack success differs by at most 0.6 points. Component studies examine the roles of difference fitting, gain-based selection, and fixed-candidate training pairs. We also derive a finite-bank calibration rule that selects episode-level policy mixtures under endpoint-risk and general-utility constraints, linking history-dependent response selection to guarantees on complete interactions.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.