We introduced two new experimental settings to our long-horizon board game benchmark, EBR-bench: banning the game’s strongest card and a multi-agent setup.
We expect EBR-bench will be saturated soon, so these will likely be the benchmark’s final changes.
Giving a model up to four subagents did little for long-horizon game scores, a data point on whether multi-agentUsing several AI agents on one problem — splitting work in parallel, checking each other, or filling different roles like planner and reviewer.Full definition → scaffolds help avoid repeated strategies. It also signals EBR-bench will saturate within months.
Terms in this piece · Glossary
benchmark — A standard public test set for comparing AI models — the shared scoreboards behind every "model X beats model Y" claim.
multi-agent — Using several AI agents on one problem — splitting work in parallel, checking each other, or filling different roles like planner and reviewer.
agent harness — The scaffolding around a model that turns it into a working agent — the loop, the tools it can call, and the rules for when to stop.
AI agent — An AI system that doesn't just answer once but works toward a goal in a loop — taking actions, reading the results, and deciding what to do next.