When Interfaces Speak: Data-Aware Generative UI Harness for Active Interaction
Abstract
Most human-agent interaction today remains text-based. However, natural language is often not a perfect medium for complex tasks, as it can lead to cognitive overload, ambiguity, information chaos, and a slow input process; ephemeral generative UIs can present structured information and guide users toward task completion. In this work, we propose **GenUI-Harness**, a multi-agent harness that pairs a Tool Agent for information retrieval and task execution with a GUI Coder Agent that identifies ambiguities and generates front-end code for structured interfaces, helping users complete their tasks. Training such a versatile coder with reinforcement learning is challenging: verifiable rewards for interactive UI generation require costly execution, while LLM-as-a-Judge rewards are prone to reward hacking. We address the first challenge with **Dynamic UX**, a lightweight package that supports dynamic interaction and reward collection within a single sandbox, and the second with **Reward Auditor**, a meta-reward mechanism that monitors reward distributions and automatically distills diagnostic patterns into a shared rubric and scoring specification. To evaluate this task, we introduce **UI-TAU Bench**, a benchmark for active human-agent interaction through generated UI code, built on **10** real-world domain databases constructed from public data sources and based on Tau-Bench tool use settings, with Lite (**300 tasks**) and Full (**1,000 tasks**) splits. Experiments show that GenUI-Harness achieves an average Pass@3 gain of **4.48 percentage points** over smolagents on Lite. Training with GenUI-Harness improves a 4B backbone from a **9.33%** baseline to **58.00%** Pass@3, even outperforming larger frontier models such as Claude Opus 5 (**46.67%**). We further show that GenUI-Harness remains robust not only on ambiguous queries, but also on non-ambiguous ones. In a reviewer survey comparing communication channels, generated UIs reduce the average number of dialogue rounds from **3.4** to **1.2**. These results show that data-aware generative interfaces can support effective task completion and reduce dialogue rounds in the evaluated database-backed workflows.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.