Vibeleaderboard
← All Intel
Intel / post

Meta previews WildArtifactBench and releases ten of its tasks

Source
x.com
Date
AI at Meta@AIatMeta
Thread · 4 parts

Today we’re also previewing WildArtifactBench, an internal evaluation framework designed to assess agents on complex, real-world tasks across diverse deliverable formats. By using win rates and Elo scores from human and agentic preference judges rather than strict ground-truth rubrics, it expands task coverage across practical multimodal workflows. We’re releasing 10 tasks from WildArtifactBench as a step forward in our ability to measure the real practical utility delivered by multimodal agents:

We’re sharing some of our design principles for WildArtifactBench:

Here are the generated artifacts of Muse Spark 1.1 and 1.2 on the Cat Mesh: https://t.co/loY5g46x1V

Here are the generated artifacts of Muse Spark 1.1 and 1.2 on the Plywood Whale: https://t.co/q5dATHjPGZ

Why it matters

A frontier lab publishing preference judged tasks gives the field an alternative to rubric bound benchmarks.

Terms in this piece · Glossary
  • AI agent — An AI system that doesn't just answer once but works toward a goal in a loop — taking actions, reading the results, and deciding what to do next.
More from AI at Meta
Recommended reads
Comments

Checking sign-in…

Loading comments…