acceptodds
Under review as a conference paper at ICLR 2027

BEHAVE: Functional Behavior Modeling Enables Self-Improving Agents for Hardware Design and Verification

Abstract

Developing agents for hardware design and verification requires reliable correctness feedback. However, a hardware specification may permit correct implementations with different latencies. Matching outputs cycle by cycle to a reference implementation can therefore reject valid designs. To address this, we introduce BEHAVE, an agentic framework for multi-turn joint hardware design and verification through functional behavior modeling. We define Behavior IR to express task functionality as executable behavior models without prescribing implementation timing beyond the specification. The agent develops a register-transfer-level (RTL) design and a behavior model as the design's verification reference. Our evaluator, BEHAVE-Sim, generates input stimuli and checks both artifacts separately against a golden behavior model, using goal-guided search to exercise additional execution cases. BEHAVE thus supports power, performance, and area (PPA) exploration across task-permitted latencies and microarchitectures. During training, the same evaluator provides verifiable reinforcement learning (RL) rewards from specification-behavior pairs without reference RTL. For self-improvement, we show that the agent can search for high-level implementations relevant to its capability gaps, construct and check specification-behavior pairs, and continue training. We release BEHAVE-Train and BEHAVE-Eval with 600 human-reviewed specification-behavior pairs for realistic hardware workloads. On BEHAVE-Eval, self-improvement from 60 seed tasks raises Qwen3.8-27B's RTL pass@1 from 55.0% to 75.0%, reaching performance comparable to RL using a 540-task pool.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.