acceptodds
Under review as a conference paper at ICLR 2027

The Last Interface: Benchmarking Language Agents on Native Terminal User Interfaces

Abstract

Terminal-oriented agent evaluation has developed quickly, but it has converged on one interaction model: the agent issues shell commands and reads streamed output. Yet much of the software that runs on servers never offers even that prompt. Editors, file managers, process monitors, debuggers, and mail clients take over the terminal, repaint one screen of characters, and make that screen their entire interface: there is no prompt to return to, no element tree to read, and no API to call. This interface class, the native text-based user interface, has gone unmeasured. We present TUI-Bench, the first execution benchmark for agents that operate native TUIs. Its 492 tasks span 77 native Linux TUI tools in 19 application categories grouped into 6 supercategories; each task starts inside the running program, exposes a plain-text character viewport with cursor coordinates, permits only keyboard actions, and is scored by deterministic verifiers that check application state, preserved data, and whether the program should still be running at the end. In a three-round evaluation of frontier models, the strongest models solve only a small fraction of scored episodes. The evaluation reveals a stable and interpretable failure pattern: agents succeed almost only where the interface has already degraded to a command line or a filter box, which suggests that they rely on pattern matching against commands rather than on continuous tracking of interface state. Whenever a task requires spanning multiple modes, holding focus, or protecting untouched state while reading, models lose control through premature exits, collateral state damage, or budget-exhausted repetition: the text on the screen is read correctly, but the interface itself is never modeled. TUI-Bench thus provides a controlled and reproducible measure of a capability gap that was previously invisible.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.