The result is a clean warning against reading scores as model-only measurements: memory, management and design can dominate the apparent capability.
“turning on two API settings we use in ChatGPT and Codex—retained reasoning and compaction—tripled scores and cut output tokens by 6x on the public task set”
OpenAI
“Benchmarks rarely measure AI models in isolation. They also measure less visible choices about API settings, harness design, and prompting.”
OpenAI
“on ARC-AGI-3, a benchmark of 2D puzzle games, GPT‑5.6 Sol scored just 7.8%, and GPT‑5.5 could barely play the games at all, scoring a paltry 0.4%”
OpenAI
“This meant that with each action, GPT‑5.6 Sol was asked to figure out the game anew, unable to remember its past thinking.”
OpenAI
“Together, retaining reasoning and compaction allow GPT‑5.6 Sol (max) to achieve roughly 3x the score with 6x fewer output tokens.”
OpenAI
Checking sign-in…
Loading comments…