BrokenArXiv
matharena.ai- Category
- Developer Tools
- Type
- TOOL
- Date
About
Tests whether models recognize false mathematical statements instead of claiming to prove them. The August 2026 edition contains 56 statements drawn from recently refuted conjectures. Native agent harnesses have coding tools but no internet. Scores normalize a revised 0–3 rubric; they are not binary problem-solving accuracy.
Why it made the leaderboard
Compare dataset edition, agent harness, tool access and grading before comparing results. A single-run observation does not establish a cross-model ranking.
Intel on BrokenArXiv
Tags
benchmarkevaluation
Media

Comments (0)
No comments yet
Editorially curated, with community endorsements as a secondary signal. Corrections welcome.