
GUI agents usually break outside the they were tuned on. This report describes a model trained across desktop, web and mobile environments with verified rewards, which is the axis that determines whether computer-use automation survives real deployment.
articleGUI-CC: Benchmarking Contextual Consistency of GUI World Models as Agent EnvironmentsLin Fu, Zheyuan Yang, Tianhui Zhang, Jinbiao Wei, Guo Gan, Boxu Liu, Yilun Zhao, Yu Rong
articleBenchmarking General Mobile Assistants in Challenging Real-World ScenariosYiqi Zhu, Feiyu Gao, Jiaxing Fan, Jiahui Zeng, Minggang Wu, Chenliang Li, Haiyang Xu, Peng Li, Ming Yan, Yang Liu
articleFramework and Benchmark for Code-Driven Agentic Testing in Web DevelopmentBin Hong, Zhenchao Zhang, Jiyuan He, Kai Zhang, Zhenya HuangChecking sign-in…
Loading comments…