← All IntelClip / AI AgentsOpus 4.7 leads, but Opus 4.8 regressed due to removed training recipe
From Vending-Bench: Long-Horizon Agent Evals — Lukas Petersson, Andon Labs · ≈2:20
“One thing that really surprised us when we ran Opus 4.8 was that it was much much worse.”
“They said that they removed a part of of the post- training recipe that was trained that that was meant to um to um do business skills.”
“Chinese models have been catching up, but it seems like um it's not by much.”
What’s in it
- Explains why Opus 4.8 scored worse than 4.7 on a business benchmark
- Ranks GLM 5.2 and GPT 5.5 against Western frontier models
- Tracks how fast Chinese AI models are closing the gap
Clip transcript
two shorter in in terms of like how longunning it is than vending bench. Um and even like two years after it was created. Um, current state-of-the-art is Opus 4.7. One thing that really surprised us when we ran Opus 4.8 was that it was much much worse. Uh, also Fable is worse. And we were like, "Oh, no, our benchmark is bad because there's something something clearly Opus 4.8 should be better than 4.7. Um, but if you look in the system card for for when Entropic released 4.8, it. They said that they removed a part of of the post- training recipe that was trained that that was meant to um to um do business skills. So, it all checked out. Um recently, GLM 5.2 has done very well and is second. Uh GP 5.5 is is third. Um and yes, um Chinese models have been catching up, but it seems like um it's not by much. They have improved a lot recently, mostly by GLM and and Kimmy. Um, but still the the frontier western ones are are uh much better. Uh, one
Comments
Sign in to comment.
Loading comments…