Understanding Brittleness to Variable Renaming in Decoder-only Code Models
Abstract
Transformer-based code models increasingly support compiler and programming- language tasks, including program analysis, performance prediction, and code generation. Recent work has shown that Transformer-based models for code are sensitive to simple semantics-preserving edits such as variable renaming. For example, a simple variable renaming can induce sizable prediction shifts in these models. The mechanisms that underlie this brittleness to variable renaming are not well understood, though. In this paper, we aim to investigate such mechanisms. In particular, we investigate where failures arise in the model, which components are involved in such failures, and whether the diversity of variable names in the training set affects the robustness of models to variable renaming. To enable our controlled experiments, we introduce TINYTRACER, a framework that generates synthetic Python programs and their execution traces under controlled training and test distributions. We use this framework to train small decoder-only Transformers from scratch to predict execution traces, while varying the diversity of variable names in the training set. We test the robustness of the trained models by taking programs that the models trace correctly and consistently replacing their variable names with other names. In these small scale models trained with low variable-name diversity, the first error in a failed trace most often occurs where a variable name should end: the model generates another identifier character instead of the required delimiter. Increasing variable-name diversity during training reduces these termination errors and improves robustness to variable names in this setting. We identify attention heads that attend to delimiters following earlier occurrences of a variable name, and find that ablating these heads reduces accuracy at name endings, supporting a causal role for these heads in identifying where variable names end. Higher training-name diversity is also associated with stronger attention to these delimiters. Experiments with another larger-scale decoder-only model (158M-parameter) trained on real Python code also show improved robustness with increased variable name diversity. Complementary analyses identify a shared difficulty in generating unfamiliar names, with errors concentrated at the initial tokens of variable names.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.