WebUIProof: Benchmarking WebUI Code Generators with UI-Agent Execution Harness
Source
arxiv.org
Author
Yun-Yun Tsai, Yuning Mao, Shiqi Wang, Junfeng Yang, Sinong Wang
Date
Why it matters
Pages that build and render often still fail interaction requirements, and static checks miss this. An AI agentAn AI system that doesn't just answer once but works toward a goal in a loop — taking actions, reading the results, and deciding what to do next.Full definition →-driven executable test harness gives a more honest measure of generated UI quality.
Key takeaways · AI-distilled
WebUIProof argues that build success and screenshots miss functional correctness, so it pairs structured specs with dense executable interaction tests covering general web UIs and 3D interactive simulations such as particle systems and physics dynamics.
Its UI-agent harnessThe scaffolding around a model that turns it into a working agent — the loop, the tools it can call, and the rules for when to stop.Full definition → runs tests in a headless browser through a plan, act, observe loop: it finds DOM elements, performs actions, watches the resulting UI and DOM changes, and checks the specified assertions.
Across eight commercial LLMs, the authors saw frequent failures on interaction requirements even when pages rendered fine, with 3D simulation interfaces the weakest area.
The same harness doubles as a reward source: training compact models such as Qwen2.5 14B and MIMO 7B with RL rewards from the executable tests improved functional completion and reduced build failures, per the authors.
Terms in this piece · Glossary
AI agent — An AI system that doesn't just answer once but works toward a goal in a loop — taking actions, reading the results, and deciding what to do next.
agent harness — The scaffolding around a model that turns it into a working agent — the loop, the tools it can call, and the rules for when to stop.