Today we’re also previewing WildArtifactBench, an internal evaluation framework designed to assess agents on complex, real-world tasks across diverse deliverable formats. By using win rates and Elo scores from human and agentic preference judges rather than strict ground-truth rubrics, it expands task coverage across practical multimodal workflows. We’re releasing 10 tasks from WildArtifactBench as a step forward in our ability to measure the real practical utility delivered by multimodal agents:

Here are the generated artifacts of Muse Spark 1.1 and 1.2 on the Plywood Whale: https://t.co/q5dATHjPGZ

A frontier lab publishing preference judged tasks gives the field an alternative to rubric bound benchmarks.
postIntroducing Muse Glimmer, an open-weight 30B-parameter model optimized for local
postIn long-horizon stress testing, Muse Code iteratively optimized GPU kernels overSign in to comment.
Loading comments…