ProtoT v2: A Simpler, More Stable, and Editable Prototype Transformer
Abstract
While state-of-the-art large language models dominate benchmarks and real-world performance, their reasoning has remained hard to trust and verify. This has inspired a new generation of language models based on interpretable components, but such models often suffer from poor scalability, efficiency, and performance. In this work, we significantly simplify and stabilize the recently published Prototype Transformer (ProtoT), leading to substantial improvements across all three of these dimensions. We also develop a custom Triton kernel for the simplified formulation that increases the speed by about , substantially surpassing all baselines. Finally, we introduce a new synthetic-prototype method that enables more flexible and powerful model editing, on par with strong baselines, while showing that our model is particularly editable. Together, these advances yield ProtoT v2, a substantially more performant, efficient, and editable prototype-based language model that retains the interpretability and robustness properties of ProtoT.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.