Veri-eRank: Verilog Data Selection via Stable Structural Adaptation Responses
Abstract
Code language models are increasingly adapted for register-transfer-level (RTL) Verilog generation from natural-language specifications. Simulation-validated corpora provide abundant executable supervision, but functional correctness does not make every RTL example equally useful to a given model: an example's adaptation value depends on the model's current capabilities. Data selection should therefore identify examples that most effectively advance the model's specification-to-design ability. Existing methods, whether loss- or gradient-based, score candidates from signals aggregated over the full sequence, without explicitly assessing whether learning from an example strengthens the model's representations of the specification-relevant RTL structures. We introduce Veri-eRank, which ranks RTL examples by changes in specification-relevant structural representations from a base model to an adapted model. Veri-eRank represents each RTL program using AST-derived structural spans, including declarations, assignments, event controls, and control flow, and weights each span by its relevance to the specification. Specifically, Veri-eRank measures the spectral entropy of the structural representations before and after adaptation, and favors examples that reduce spectral entropy after adaptation. To avoid scores driven by a particular local observation, Veri-eRank evaluates resampled structural views and favors responses that are both strong and consistent. Theoretically, our analysis shows that preserving structural directions while reducing isotropic variation lowers spectral entropy after adaptation, motivating our spectral-response score. Under a 20% data budget, Veri-eRank attains the highest mean Pass@1 among the compared data selection methods in all six settings across two models and three RTL benchmarks. Five-seed macro-average Pass@1 reaches for Qwen2.5-Coder and for Seed-Coder, exceeding SPICE by 5.42 and 2.24 percentage points and Full SFT by 1.05 and 1.98 points, respectively.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.