Vibeleaderboard
Index / tool
Visit arcprize.org
Category
Developer Tools
Type
TOOL
GitHub
65 stars
Added
Jul 31, 2026

About

Measures skill acquisition and generalization on novel abstract tasks, including interactive agentic environments in ARC-AGI-3.

Why it made the leaderboard

Benchmark results are decision inputs for choosing models and agent harnesses; follow the methodology and current leaderboard, not a single headline score.

Tags

benchmarkreasoninggeneralizationagentic intelligence

Tech Stack

Python

Comments (0)

No comments yet

Indexed by a proprietary survey. Corrections welcome.