Does It Render Everywhere? A Study of Cross-Environment Compatibility in MLLM-Generated Webpages
- Source
- Ziyun Guo, Jingyu Xiao, Yuqiang Sun, Yintong Huo
- Author
- Ziyun Guo, Jingyu Xiao, Yuqiang Sun, Yintong Huo
- Date

- eval — A repeatable test for AI quality — a set of tasks plus scoring — used the way software teams use test suites, because model output is too variable to judge by eyeballing.
Puts a number on a failure mode most AI frontend misses entirely, since visual fidelity is usually scored in a single fixed browser.
“Multimodal Large Language Models (MLLMs) have been increasingly adopted to automate webpage generation from visual designs (e.g., screenshots). However, existing evaluations are limited to visual fidelity assessment under a fixed browser-device configuration. Such a setting overlooks the cross-environment rendering compatibility for real-world deployments.”
Ziyun Guo, Jingyu Xiao, Yuqiang Sun, Yintong Huo
“Specifically, we construct WebCompat, a dataset of 2,032 annotated instances, comprising webpages generated by 8 representative AI tools, each rendered across 9 browser-and-device combinations.”
Ziyun Guo, Jingyu Xiao, Yuqiang Sun, Yintong Huo
“Our findings reveal that 68% of generated webpages exhibit at least one compatibility issue, underscoring the pervasive reliability concerns surrounding MLLM-generated front-end artifacts.”
Ziyun Guo, Jingyu Xiao, Yuqiang Sun, Yintong Huo
“The most prevalent symptoms are failures that disrupt the entire page layout (88.3%): pages shrink directly to fit the target screen with too small fonts, or exhibit scale mismatches that produce cut-off content.”
Ziyun Guo, Jingyu Xiao, Yuqiang Sun, Yintong Huo
“Guided by the findings, we develop XCompat, a lightweight offline compatibility issue detector that combines visual screenshots and the structural DOM tree for analysis. It achieves an F1 score of 0.903 on the WebCompat-test, outperforming the existing compatibility checking tools and LLM baselines.”
Ziyun Guo, Jingyu Xiao, Yuqiang Sun, Yintong Huo
articleGUI-CC: Benchmarking Contextual Consistency of GUI World Models as Agent EnvironmentsLin Fu, Zheyuan Yang, Tianhui Zhang, Jinbiao Wei, Guo Gan, Boxu Liu, Yilun Zhao, Yu Rong
articleFramework and Benchmark for Code-Driven Agentic Testing in Web DevelopmentBin Hong, Zhenchao Zhang, Jiyuan He, Kai Zhang, Zhenya Huang
articleCompliance, Capability, and Conflict: Benchmarking Multimodal LLMs under System MessagesJuan Yeo, Geewook Kim
Checking sign-in…
Loading comments…