acceptodds
Under review as a conference paper at ICLR 2027

What Accuracy Hides In Learned Transformer Programs

Abstract

Transformer Programs (TP) (Friedman et al., 2023) constrain a Transformer’s architecture so that its learned computation can be extracted as readable Python code. TP expose learned computations, but token accuracy can hide incorrect complete answers. In this work, we examine how program inspection provides insights beyond accuracy scores and helps guide training. We first evaluated TP on nine algorithmic tasks. We compared token accuracy with accuracy on complete output sequences and checked for failures missed by sampled tests. Further, we assessed training data improvements informed by program inspection and failure analysis, finding gains in sequence accuracy on several tasks. We examined a counting task where a program achieved 100% accuracy on sampled test sequences but failed on another input using the symbols and lengths allowed during training. The program counted correctly, but the required answer was missing from its output labels. After adding training examples with larger counts, the retrained models passed exhaustive evaluation over the bounded domain. In the frequency ranking task, inspection showed how correct counts contributed to incorrect answers in programs trained before and after changes to the training data. Finally, we studied constraints imposed by TP training and operations. The TP training process can lead to accuracy losses late in training. Our reference constructions also show how composing the restricted TP operations can use many layers or large lookup tables.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.