SIGMA: Self-Improving Alignment Generalization from a Model Spec
Abstract
LLM agents are increasingly capable of executing complex tasks and of recursively improving themselves on easy-to-verify objectives such as software engineering and mathematics. Since alignment is much harder to verify, this creates a growing risk of capabilities increasing without appropriate safety alignment: models may act dangerously in consequential situations their safety training never anticipated, especially as their capabilities expand to auto-research and cybersecurity domains. Existing approaches often focus on capability self-improvement using verifiable feedback, or on alignment training with supervision from stronger models or curated data, creating an external supervision bottleneck for alignment. To address this bottleneck, we ask whether current models can improve their own safety alignment. We propose SIGMA, a data generation and training pipeline enabling alignment self-improvement that generalizes to out-of-distribution settings. Given only a Model Spec that states the model's desired behavior and high-level prompts defining its goal, SIGMA leverages a model's general reasoning capabilities to strengthen its own safety reasoning capability. SIGMA first performs spec-guided task synthesis, using the candidate model as a task designer agent to generate diverse alignment dilemma scenarios and convert them into high-quality alignment training tasks that stress-test the model’s understanding of the Model Spec. Next, SIGMA conducts self-judged alignment training through supervised fine-tuning and rubric-based reinforcement learning with the model itself as the reward model. Despite training only on single-turn chat data describing alignment dilemmas, SIGMA improves safety alignment in multi-turn agentic environments (AgentHarm harmfulness decreases from 22.6 to 14.8; Agentic Misalignment decreases from 79.1 to 3.8), outperforms Deliberative Alignment and Constitutional AI baselines, and retains general capability. Analyses show that a Model Spec that balances harmlessness and helpfulness, test-time reasoning for safety deliberation, and high-quality rubrics produced by SIGMA's task designer agent are crucial for effective self-improvement.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.