When Interfaces Speak: Data-Aware Generative UI Harness for Active Interaction
Quick Answer
This paper shows that GenUI-Harness combines a Tool Agent and a GUI Coder Agent to enhance human-agent interactions, achieving a 4.48% Pass@3 improvement over smolagents on the Lite benchmark.
Quick Take
Training with this system boosts a 4B model's performance from 9.33% to 58.00% Pass@3, outperforming Claude Opus 5. The approach significantly reduces dialogue rounds from 3.4 to 1.2, demonstrating the effectiveness of data-aware generative interfaces.
Key Points
- GenUI-Harness utilizes Dynamic UX for interactive UI generation and reward collection.
- Reward Auditor monitors and refines reward distributions for effective training.
- UI-TAU Bench includes 10 real-world databases for evaluating human-agent interactions.
- The system shows robustness against both ambiguous and clear queries.
- Generated UIs streamline task completion, reducing dialogue rounds significantly.
DeepSignal Analysis
What happened
The GenUI-Harness system integrates a Tool Agent and a GUI Coder Agent to improve human-agent interactions. It achieved a 4.48% Pass@3 improvement over smolagents on the Lite benchmark and significantly enhanced a 4B model's performance from 9.33% to 58.00% Pass@3, surpassing Claude Opus 5.
Key evidence
- GenUI-Harness combines a Tool Agent for task execution with a GUI Coder Agent for generating structured interfaces, addressing challenges in interactive UI generation.
- The system's training improved a 4B model's Pass@3 score from 9.33% to 58.00%, outperforming Claude Opus 5, which scored 46.67%.
- Dialogue rounds were reduced from an average of 3.4 to 1.2 when using generated UIs, indicating more efficient task completion.
Why it matters
This development highlights a shift towards more efficient human-agent interactions by utilizing data-aware generative interfaces. The significant performance improvements suggest potential applications in various domains where complex tasks are common. Reducing dialogue rounds can enhance user experience and productivity, making this technology relevant for future AI applications.
What to watch
Paper Resources
📖 Reader Mode
~2 min readAbstract:Most human-agent interaction today remains text-based. Natural language can impose cognitive overload, ambiguity, information chaos, and slow input for complex tasks; ephemeral generative UIs can present structured information and guide users toward task completion. We propose GenUI-Harness, a multi-agent harness pairing a Tool Agent for information retrieval and task execution with a GUI Coder Agent that identifies ambiguities and generates front-end code for structured interfaces. Training the coder with reinforcement learning is challenging: verifiable rewards for interactive UI generation require costly execution, while LLM-as-a-Judge rewards are prone to reward hacking. We address the first challenge with Dynamic UX, a lightweight package for dynamic interaction and reward collection in a single sandbox, and the second with Reward Auditor, a meta-reward mechanism that monitors reward distributions and distills diagnostic patterns into a shared rubric and scoring specification. We introduce UI-TAU Bench, a benchmark for active human-agent interaction through generated UI code, built on 10 real-world domain databases constructed from public data sources and based on Tau-Bench tool-use settings, with Lite (300 tasks) and Full (1,000 tasks) splits. GenUI-Harness achieves an average Pass@3 gain of 4.48 percentage points over smolagents on Lite. Training with GenUI-Harness improves a 4B backbone from 9.33% to 58.00% Pass@3, outperforming larger frontier models such as Claude Opus 5 (46.67%). GenUI-Harness also remains robust on ambiguous and non-ambiguous queries. In a reviewer survey comparing communication channels, generated UIs reduce average dialogue rounds from 3.4 to 1.2. These results show that data-aware generative interfaces can support effective task completion and reduce dialogue rounds in evaluated database-backed workflows.
| Comments: | 35 pages, 7 figures. Code: this https URL |
| Subjects: | Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2610.11123 [cs.AI] |
| (or arXiv:2610.11123v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2610.11123 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Xiaolong Li [view email]
[v1]
Thu, 8 Oct 2026 02:49:35 UTC (979 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.AI
See more →HOBA: Hierarchical On-Policy Bidding Agents for Adaptive Online Advertising
HOBA (Hierarchical On-policy Bidding Agents) is a novel hierarchical reinforcement learning framework that enhances online advertising bidding systems by improving adaptability and reducing hyperparameter tuning costs. It utilizes a for hyperparameter inference, a SARSA agent for expert model selection, and a dynamic expert pool for bid execution, achieving a +3.6% increase in target cost during large-scale deployment and outperforming state-of-the-art baselines on AuctionNet.