RealControlBench: A Comprehensive Benchmark for Fine-Grained LLM and LLM Agent steer in the Wild
Abstract
Large Language Models (LLMs) are increasingly deployed in settings that require precise control over their behavior, from safety and style to reasoning and agentic decision-making. Steering methods offer an attractive alternative to prompting or retraining by enabling targeted and inexpensive interventions; yet, their effectiveness is difficult to assess systematically. However, existing steering benchmarks are often built around simplified and isolated generation tasks, evaluating a single behavioral attribute under controlled settings that do not reflect the diversity, interaction complexity, and capability-preservation requirements of practical deployment. As a result, strong performance on existing benchmarks may provide only a limited picture of whether a steering method remains reliable in realistic settings. We therefore introduce RealControlBench, a comprehensive benchmark for evaluating fine-grained steering of LLMs and LLM agents across diverse behavioral targets and interaction settings. RealControlBench extends steering evaluation beyond isolated generations to complex, execution-grounded, and multi-turn agent interactions, covering five complementary dimensions: concept and style control, instruction-following control, safety steering, reasoning-performance and chain-of-thought length control, and agentic decision control. Rather than measuring steering success alone, the benchmark jointly evaluates target controllability and capability preservation, including instruction following, reasoning performance, and downstream agent utility and safety. It further examines whether steering effects remain specific and robust across prompt variations, domains, and model architectures. Our evaluation shows that existing steering methods remain unreliable across realistic control objectives: they often fail to consistently induce the intended behavior and can substantially degrade the model's underlying task performance. These results suggest that strong steering effects in controlled settings do not necessarily translate into reliable control in more realistic evaluations. By providing a unified, model-agnostic steering evaluation framework, RealControlBench aims to support the development of steering methods that are not only effective, but also precise, robust, and useful in realistic AI deployments.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.