Vibeleaderboard
Index / article

Two settings tripled GPT-5.6 Sol’s ARC-AGI-3 score

openai.com
Visit openai.com
Category
Other
Type
ARTICLE
Added
Jul 30, 2026

About

OpenAI found that retaining reasoning and replacing rolling truncation with compaction moved GPT-5.6 Sol from 13.3% to 38.3% on the ARC-AGI-3 public set while using six times fewer output tokens.

Why it made the leaderboard

The result is a clean warning against reading benchmark scores as model-only measurements: memory, context management and harness design can dominate the apparent capability.

Comments (0)

No comments yet

Indexed by a proprietary survey. Corrections welcome.