Two settings tripled GPT-5.6 Sol’s ARC-AGI-3 score
openai.com- Category
- Other
- Type
- ARTICLE
- Added
- Jul 30, 2026
About
OpenAI found that retaining reasoning and replacing rolling truncation with compaction moved GPT-5.6 Sol from 13.3% to 38.3% on the ARC-AGI-3 public set while using six times fewer output tokens.
Why it made the leaderboard
The result is a clean warning against reading benchmark scores as model-only measurements: memory, context management and harness design can dominate the apparent capability.
Comments (0)
No comments yet
Indexed by a proprietary survey. Corrections welcome.