Vibeleaderboard
Index / tool

BrokenArXiv

matharena.ai
Visit matharena.ai
Category
Developer Tools
Type
TOOL
Date

About

Tests whether models recognize false mathematical statements instead of claiming to prove them. The August 2026 edition contains 56 statements drawn from recently refuted conjectures. Native agent harnesses have coding tools but no internet. Scores normalize a revised 0–3 rubric; they are not binary problem-solving accuracy.

Why it made the leaderboard

Compare dataset edition, agent harness, tool access and grading before comparing results. A single-run observation does not establish a cross-model ranking.

Intel on BrokenArXiv

More in Intel

Tags

benchmarkevaluation

Media

BrokenArXiv

Comments (0)

No comments yet

Editorially curated, with community endorsements as a secondary signal. Corrections welcome.