
To understand whether we're making genuine progress on reasoning, we entered our AI models in five STEM Olympiad competitions. The results: 🏅 Asian Physics Olympiad (APhO): Perfect score, theory exam 🏅 International Physics Olympiad (IPhO): Perfect score, theory exam 🥇 International Mathematical Olympiad (IMO): Gold medal 🥇 International Chemistry Olympiad (IChO): Gold-medal-level performance 🥇 Romanian Masters of Mathematics (RMM): Gold-medal-level performance The types of problems in the Olympiad competitions are exceptionally hard, demanding deep chains of reasoning, creative insight, and flawless argumentation. To test pure reasoning capability, we disallowed all tool use, meaning no search, no coding, and no calculator. We have deep admiration for the contestants and committees behind these competitions, and are grateful for their support in enabling our participation.
Olympiad results under competition conditions give practitioners an independent read on how far frontier reasoning has actually moved.
Checking sign-in…
Loading comments…