AI Alignment via Incentives and Correction
Abstract
We study AI alignment through law-and-economics models of deterrence, in which misconduct is a strategic response to the probability of detection and the severity of punishment. The same logic arises in agentic AI pipelines: a solver may gain from a persuasive but incorrect answer, while an auditor must decide whether costly monitoring is worthwhile. This creates a coupled incentive problem: stronger penalties deter the solver but can weaken the auditor's incentive to inspect, since once the solver seems reliable, auditing mostly incurs cost and rarely catches errors. This perspective also changes what should count as a post-training signal: rather than rewarding the final answer alone, a solver–auditor pipeline lets rewards target the full correction event. We formalize this in a two-agent model in which a principal chooses these rewards, making reward design a bilevel optimization problem judged by the behavior it induces. We propose a bandit-based outer loop that searches over reward profiles using noisy interaction feedback. On an LLM coding pipeline, adaptive reward profiles maintain oversight pressure and improve held-out principal value over static hand-designed rewards, reducing hallucinated incorrect attempts by 52%.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.