Vibeleaderboard
← All Intel
Intel / post

Cua Adds a Guide for Running OSWorld Benchmarks on Its Fleets

Source
Cua
Date
Cua@trycua
Thread · 4 parts

1/ You can now run OSWorld tasks on Cua Cloud Fleets. Our new guide shows how to turn the official OSWorld Ubuntu image into a Fleet desktop with Cua Driver, then run your agent and evaluate the tasks. Get started: https://t.co/T3Xx7ZcasW

2/ Bring the benchmark desktop with you. Start with the official OSWorld image and its applications. Add Cua Driver, publish the VM disk to a registry, and claim a desktop through the Sandbox SDK.

3/ Set up the task. Run your agent. Evaluate the result. Use OSWorld's setup tools and evaluators against the Fleet desktop, with Cua Driver for agent actions. Claim a fresh desktop for the next task; Fleet doesn't revert snapshots.

4/ Ready to try it? The guide includes the image recipe and a Python sample for claiming the desktop, taking a screenshot, and calling Cua Driver. You'll need Fleet credentials, a registry image, and cua-sandbox from main. Docs: https://t.co/G7Lt9LKrXL Start a Fleet:

Key takeaways · AI-distilled
  • The recipe starts from the official OSWorld Ubuntu image and its applications, adds Cua Driver, publishes the VM disk to a registry, and claims a desktop through the SDK.
  • Task setup and scoring use OSWorld's own setup tools and evaluators against the Fleet desktop, with Cua Driver handling actions. Because Fleet does not revert snapshots, each task needs a freshly claimed desktop.
  • The guide includes a Python sample that claims the desktop, takes a screenshot and calls Cua Driver. It requires Fleet credentials, a registry image and cua-sandbox installed from main.
Terms in this piece · Glossary
  • benchmarkA standard public test set for comparing AI models — the shared scoreboards behind every "model X beats model Y" claim.
  • AI agentAn AI system that doesn't just answer once but works toward a goal in a loop — taking actions, reading the results, and deciding what to do next.
  • sandboxAn isolated environment where AI-generated code or agent actions run without being able to touch anything real.
Why it matters

Lowers the setup cost for running the standard OSWorld computer-use against agents built on Cua's stack.

More from Cua
Recommended reads
Comments

Checking sign-in…

Loading comments…