← All IntelClip / AI AgentsARC-AGI-3 blown open in hours, and the benchmark politics that followed
From Recursive Coding Agents - Raymond Weitekamp, OpenProse · ≈7:16
Documents both a striking harness-over-model result and the resulting evaluation dispute — leaderboards designed around no-tool-calling now need separate open-harness tracks to score RLM-style systems at all.
What’s in it
- Documents both a striking harness-over-model result and the resulting evaluation dispute — leaderboards designed around no-tool-calling now need separate open-harness tracks to score RLM-style systems at all.
Clip transcript
powerful. So powerful that they are arguably too hot to benchmark. So, two examples here, on the left, a very high-profile uh case where um the Symbolica team has this RLM agent harness called Agentica. Within hours of Arc AGI-3 being released, where the top scores of all the frontier models were around two or three percent, the uh Symbolica team showed 30-something percent. This is crazy. They blew it out of the water within hours using RLMs as a framework. So much so that is uh, very much upset the Arc Prize team. And so, uh, they gave them what I'm interpreting as a consolation tweet. Uh, which as far as I'm reading the situation was essentially saying, you know, la-di-da, congratulations, um, but you didn't solve the problem the right way. And we don't like RLM harnesses. And so, uh, you can have this nice tweet, but we refuse to actually do the full private part of the ArcadeGI evaluation. Uh, which to me is just insane.
Comments
Checking sign-in…
Loading comments…