Vibeleaderboard
← All Intel
Intel / post

Vals AI: higher reasoning effort has diminishing returns on math proofs

Source
Vals AI
Date
Vals AI@ValsAI

How does "thinking harder" change how models perform on math proofs? We ran GPT-6 Astra, Opus 5, and Opus 5.5 on every reasoning level: low, medium, high, xhigh, and max on our benchmark, Proof Bench v1.1.

Why it matters

Vals AI's Proof Bench shows Opus 5.5 at medium effort hits 99% accuracy for a fraction of the cost of pushing to max, while some models plateau regardless of added reasoning effort.

More from Vals AI
Recommended reads
Comments

Checking sign-in…

Loading comments…