Vibeleaderboard
Index / agent
Visit github.com
Category
AI Agents
Rank
No. 1054Tools index
Pricing
Open Source
Type
AGENT
Use case
Research & Education
Interfaces
CLI
Builder
@alxfazio
GitHub
310 stars
Date

About

ARCgentica is an agentic AI system that solves ARC-AGI-2 challenges by deploying LLM-powered sub-agents to analyze input-output grid examples, write Python transform programs, and evaluate them against test inputs. It achieved 85.28% on the ARC-AGI-2 public evaluation using Claude Opus 4.6, making it one of the top-performing open solutions to this benchmark. It's useful for researchers and developers studying AI reasoning, program synthesis, and abstract pattern recognition.

What it does

A research harness for the ARC-AGI puzzles, where each task shows a few colored grids before and after some hidden rule and asks for that rule. A lead model works inside a live Python session, pokes at the grids with numpy and scipy, and can hand hypotheses to helper models, which can hand work down again. Its final answer is a Python function rather than a grid. The harness runs that function on the unseen puzzles in a separate process with a five second limit, keeps the first two usable answers per puzzle, and grades them against the official answer key.

Why it's ranked here

The result is the reason: the authors report 85.28% on the ARC-AGI-2 public evaluation set with Claude Opus 4.6 at $6.94 per task, and they committed the logs of that run plus a command to re-grade it. The code is small, MIT licensed and pinned to exact dependency versions, so the claim can be audited rather than taken on trust. What limits it is scope: it solves one benchmark and nothing else, and reproducing the headline number takes real money and roughly 25 minutes per task on average.

What's good

The design is short enough to read in an afternoon. Grading is exact match on the whole grid, with a softer cell-by-cell measure the models use while they iterate, and the final tally follows the competition rule of two guesses per test input. Answers are code, so every prediction can be rerun and inspected. Each run writes its settings, every attempt's result and every agent's transcript to disk, and an interrupted run resumes by ID, skipping puzzles already done. Delegation is capped at ten agents per attempt, counted across the whole tree, so a runaway chain of helpers hits a hard wall.

Tradeoffs

The recommended server launch turns sandboxing off, and generated programs run as ordinary child processes with your user's permissions, so model-written code executes on your machine. Retries fire only on one specific OpenAI content-policy error; other failures are recorded and not retried, although the documentation describes retries for transient errors. The documented default model differs from the one the code actually defaults to, and the server launch command is missing a line continuation, so the port flag never reaches the server. Python is pinned to exactly 3.12.11. The reported run used high reasoning effort, not the extra-high default. Continuous integration runs formatting and lint checks only.

How to use it well

Treat it as a reference implementation if you study program synthesis or multi-agent orchestration, or as the baseline to beat on ARC. Start by re-grading the committed run, which costs nothing, then try ten puzzles before a full pass: at the reported averages each task costs about seven dollars and takes about 25 minutes. Run the agent server in a container or virtual machine, since sandboxing is off, and fix the broken line in its launch command. It is not a general agent framework; for that, look at the Agentica SDK it is built on.

Technical notes+

main.py builds Problem objects from the data/arc-prize-<year> evaluation JSON and runs attempt_problem under an asyncio.Semaphore (default 60). solve.py runs num_attempts attempts per problem in parallel; each constructs the Agent in arc_agent/agent.py, whose call_agent spawns an Agentica agent (cache_ttl 1h) with accuracy, soft_accuracy, example_to_diagram, the grid dataclasses and call_agent itself in REPL scope, so sub-agents recurse through one shared agents list capped at max_num_agents. The returned FinalSolution transform_code is written to a temp file and run through asyncio.create_subprocess_exec with the environment reduced to PYTHONHASHSEED and a timeout (default 5 seconds). The retry branch in _attempt triggers only when the error string contains OPENAI_ERROR. score.py takes the first two non-empty outputs per test input (pass_at_two). main.py defaults --model to anthropic/claude-opus-4-6 while README.md lists openai/gpt-5.2 as default, and README.md's server command lacks a trailing backslash after --max-concurrent-invocations 1200. The committed final config.json records reasoning_effort high. .github/workflows/ci.yml runs ruff format --check and ruff check via Nix, with no test step.

Observed

License
MIT (LICENSE)
Language
Python, pinned to exactly 3.12.11 in pyproject.toml
Packaging
uv-managed project run from a clone; pyproject.toml declares no console entry point
Interface
Command-line runner with argparse flags (main.py); requires a separately cloned Agentica server
Model providers
OpenAI, Anthropic or OpenRouter through the Agentica server
Selectable models
Three hardcoded choices: openai/gpt-5.2, anthropic/claude-opus-4-5, anthropic/claude-opus-4-6
Code execution
Generated programs run in a plain subprocess with a 5 second default timeout; README.md server command sets sandbox mode to no_sandbox
Scoring
Exact grid match on the first two non-empty outputs per test input (score.py)
Reported run
Logs and configuration of the reported run are committed (output/2025/anthropic/claude-opus-4-6/final/config.json)
Continuous integration
ruff format and ruff check only, no test step (.github/workflows/ci.yml)

Read from README.md, pyproject.toml, LICENSE, main.py, solve.py, score.py, arc_agent/agent.py, arc_agent/prompts.py, arc_agent/types.py, common.py, output/2025/anthropic/claude-opus-4-6/final/config.json, .github/workflows/ci.yml.

What it can do

  • Solve ARC-AGI grid puzzles autonomously

    ARC-AGI input-output grid examples → Predicted output grids for test inputs

  • Synthesize Python transformation programs

    Grid pattern examples showing input-output relationships → Executable Python code that transforms input grids to output grids

  • Analyze abstract visual patterns in grids

    ARC-AGI puzzle grid data → Identified transformation rules and pattern descriptions

  • Evaluate generated programs against test inputs

    Python transform programs and test grid inputs → Evaluated solutions with pass/fail scoring results

  • Run multi-agent reasoning pipelines

    ARC-AGI puzzle and selected LLM backend (OpenAI, Anthropic, or OpenRouter) → Coordinated sub-agent analysis and solution attempts

  • Score and benchmark solution performance

    Completed puzzle run results → Accuracy scores and performance metrics across puzzle sets

  • Generate detailed run logs for AI reasoning research

    Executed ARC-AGI solving sessions → Full logs of agent reasoning steps, program attempts, and evaluations

Tags

arc-agillmagentic-aiprogram-synthesisbenchmarkingpythonclaudeabstract-reasoning

Tech Stack

Python

Comments (0)

No comments yet

Editorially curated, with community endorsements as a secondary signal. Corrections welcome.