acceptodds
Under review as a conference paper at ICLR 2027

Omni TM-AE: A Scalable and Interpretable Embedding Model Using the Full Tsetlin Machine State Space

Abstract

The increasing complexity of large-scale language models has amplified concerns regarding their interpretability and reusability. While traditional embedding models like Word2Vec and GloVe offer scalability, they lack transparency and often behave as black boxes. Conversely, interpretable models such as the Tsetlin Machine (TM) have shown promise in constructing explainable learning systems, though they previously faced limitations in scalability and reusability. In this paper, we introduce the Omni Tsetlin Machine Autoencoder (Omni TM-AE), a novel embedding model that fully exploits the information contained in the TM's state matrix, including literals previously excluded from clause formation. This method enables the construction of reusable, interpretable embeddings through a single training phase. We further introduce the Tsetlin State-Space Explorer (TSSE), an interactive interpretability tool for analyzing clause-level state distributions and tracing literal-state changes through training, linking learned representations to clauses, feedback events, and supporting training evidence. Extensive experiments across semantic similarity, sentiment classification, word analogy, document clustering, and quantitative interpretability evaluations show that Omni TM-AE performs competitively with mainstream embedding models while providing vocabulary-level traceability. These results demonstrate that it is possible to balance performance, scalability, and interpretability in Natural Language Processing (NLP) systems without resorting to opaque architectures.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.