
What tasks are left that humans find easy but today's models still find hard? Two such tasks are computer use and games. We’re launching CUA-Bench, a benchmark testing how well AI can use a keyboard and mouse across 6 games (3 kept private) in real time. To saturate it, models will need to output real-time actions and learn continuously from video, not just text.
A new measuring real-time keyboard/mouse control across games highlights that agents still can't act continuously from video, a concrete capability gap for anyone building computer-use agents.
Checking sign-in…
Loading comments…