Two settings tripled GPT-5.6 Sol’s ARC-AGI-3 score
Source
OpenAI
Author
OpenAI
Date
Why it matters
The result is a clean warning against reading benchmark scores as model-only measurements: memory, context management and harness design can dominate the apparent capability.
OpenAI found that retaining reasoning and replacing rolling truncation with compaction moved GPT-5.6 Sol from 13.3% to 38.3% on the ARC-AGI-3 public set while using six times fewer output tokens.
Transcript
OpenAI re-ran ARC-AGI-3 through the Responses API with retained reasoning and compaction. The same GPT-5.6 Sol model moved from 13.3% to 38.3% on the public set while producing six times fewer output tokens.