OriGent: Post-Training Hardware Engineering Agents through Scalable Environment Synthesis
Abstract
LLM-based coding agents have made substantial progress in solving software engineering tasks. Applying these agents to hardware engineering requires them to reason about circuit behavior, use simulation and electronic design automation (EDA) tools, and coordinate changes across register-transfer-level (RTL) designs, configuration files, and verification code. Learning these workflows requires hardware-specific interaction experience, but much of the relevant development data is proprietary, limiting public training resources. We introduce OriGent, a framework that constructs executable training environments from the commit history of public hardware repositories. Starting from merged pull requests, OriGent reconstructs repository states, builds target tests validated to fail before the reference repair and pass afterward, and generates problem statements aligned with the tested behavior. Separating reusable environment setup from task-specific construction yields over 18,000 executable tasks from 90 repositories spanning RTL design, verification, EDA software, hardware domain-specific languages, and firmware. We post-train Qwen3.6-35B-A3B through supervised fine-tuning on successful teacher trajectories collected in these environments, followed by online reinforcement learning. On HWE-Bench, a repository-level hardware engineering benchmark, the resolve rate increases from 39.3% for the base model to 55.9% after supervised fine-tuning and 64.3% after reinforcement learning. OriGent outperforms KAT-Coder-V2.5-Dev (51.3%), which is post-trained for agentic coding from the same base model. On a 200-task SWE-bench Pro subset, OriGent improves resolve rate from 44.0% to 58.0%, showing transfer from hardware-focused training to software engineering.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.