
ARCgentica
github.com/symbolica-ai/arcgentica- Category
- AI Agents
- Rank
- No. 1054Tools index
- Pricing
- Open Source
- Type
- AGENT
- Use case
- Research & Education
- Interfaces
- CLI
- Builder
- @alxfazio
- GitHub
- 310 stars
- Date
About
ARCgentica is an agentic AI system that solves ARC-AGI-2 challenges by deploying LLM-powered sub-agents to analyze input-output grid examples, write Python transform programs, and evaluate them against test inputs. It achieved 85.28% on the ARC-AGI-2 public evaluation using Claude Opus 4.6, making it one of the top-performing open solutions to this benchmark. It's useful for researchers and developers studying AI reasoning, program synthesis, and abstract pattern recognition.
What it does
A research harness for the ARC-AGI puzzles, where each task shows a few colored grids before and after some hidden rule and asks for that rule. A lead model works inside a live Python session, pokes at the grids with numpy and scipy, and can hand hypotheses to helper models, which can hand work down again. Its final answer is a Python function rather than a grid. The harness runs that function on the unseen puzzles in a separate process with a five second limit, keeps the first two usable answers per puzzle, and grades them against the official answer key.
Why it's ranked here
The result is the reason: the authors report 85.28% on the ARC-AGI-2 public evaluation set with Claude Opus 4.6 at $6.94 per task, and they committed the logs of that run plus a command to re-grade it. The code is small, MIT licensed and pinned to exact dependency versions, so the claim can be audited rather than taken on trust. What limits it is scope: it solves one benchmark and nothing else, and reproducing the headline number takes real money and roughly 25 minutes per task on average.
What's good
The design is short enough to read in an afternoon. Grading is exact match on the whole grid, with a softer cell-by-cell measure the models use while they iterate, and the final tally follows the competition rule of two guesses per test input. Answers are code, so every prediction can be rerun and inspected. Each run writes its settings, every attempt's result and every agent's transcript to disk, and an interrupted run resumes by ID, skipping puzzles already done. Delegation is capped at ten agents per attempt, counted across the whole tree, so a runaway chain of helpers hits a hard wall.
Tradeoffs
The recommended server launch turns sandboxing off, and generated programs run as ordinary child processes with your user's permissions, so model-written code executes on your machine. Retries fire only on one specific OpenAI content-policy error; other failures are recorded and not retried, although the documentation describes retries for transient errors. The documented default model differs from the one the code actually defaults to, and the server launch command is missing a line continuation, so the port flag never reaches the server. Python is pinned to exactly 3.12.11. The reported run used high reasoning effort, not the extra-high default. Continuous integration runs formatting and lint checks only.
How to use it well
Treat it as a reference implementation if you study program synthesis or multi-agent orchestration, or as the baseline to beat on ARC. Start by re-grading the committed run, which costs nothing, then try ten puzzles before a full pass: at the reported averages each task costs about seven dollars and takes about 25 minutes. Run the agent server in a container or virtual machine, since sandboxing is off, and fix the broken line in its launch command. It is not a general agent framework; for that, look at the Agentica SDK it is built on.
Technical notes+
main.py builds Problem objects from the data/arc-prize-<year> evaluation JSON and runs attempt_problem under an asyncio.Semaphore (default 60). solve.py runs num_attempts attempts per problem in parallel; each constructs the Agent in arc_agent/agent.py, whose call_agent spawns an Agentica agent (cache_ttl 1h) with accuracy, soft_accuracy, example_to_diagram, the grid dataclasses and call_agent itself in REPL scope, so sub-agents recurse through one shared agents list capped at max_num_agents. The returned FinalSolution transform_code is written to a temp file and run through asyncio.create_subprocess_exec with the environment reduced to PYTHONHASHSEED and a timeout (default 5 seconds). The retry branch in _attempt triggers only when the error string contains OPENAI_ERROR. score.py takes the first two non-empty outputs per test input (pass_at_two). main.py defaults --model to anthropic/claude-opus-4-6 while README.md lists openai/gpt-5.2 as default, and README.md's server command lacks a trailing backslash after --max-concurrent-invocations 1200. The committed final config.json records reasoning_effort high. .github/workflows/ci.yml runs ruff format --check and ruff check via Nix, with no test step.
Observed
- License
- MIT (LICENSE)
- Language
- Python, pinned to exactly 3.12.11 in pyproject.toml
- Packaging
- uv-managed project run from a clone; pyproject.toml declares no console entry point
- Interface
- Command-line runner with argparse flags (main.py); requires a separately cloned Agentica server
- Model providers
- OpenAI, Anthropic or OpenRouter through the Agentica server
- Selectable models
- Three hardcoded choices: openai/gpt-5.2, anthropic/claude-opus-4-5, anthropic/claude-opus-4-6
- Code execution
- Generated programs run in a plain subprocess with a 5 second default timeout; README.md server command sets sandbox mode to no_sandbox
- Scoring
- Exact grid match on the first two non-empty outputs per test input (score.py)
- Reported run
- Logs and configuration of the reported run are committed (output/2025/anthropic/claude-opus-4-6/final/config.json)
- Continuous integration
- ruff format and ruff check only, no test step (.github/workflows/ci.yml)
Read from README.md, pyproject.toml, LICENSE, main.py, solve.py, score.py, arc_agent/agent.py, arc_agent/prompts.py, arc_agent/types.py, common.py, output/2025/anthropic/claude-opus-4-6/final/config.json, .github/workflows/ci.yml.
What it can do
Solve ARC-AGI grid puzzles autonomously
ARC-AGI input-output grid examples → Predicted output grids for test inputs
Synthesize Python transformation programs
Grid pattern examples showing input-output relationships → Executable Python code that transforms input grids to output grids
Analyze abstract visual patterns in grids
ARC-AGI puzzle grid data → Identified transformation rules and pattern descriptions
Evaluate generated programs against test inputs
Python transform programs and test grid inputs → Evaluated solutions with pass/fail scoring results
Run multi-agent reasoning pipelines
ARC-AGI puzzle and selected LLM backend (OpenAI, Anthropic, or OpenRouter) → Coordinated sub-agent analysis and solution attempts
Score and benchmark solution performance
Completed puzzle run results → Accuracy scores and performance metrics across puzzle sets
Generate detailed run logs for AI reasoning research
Executed ARC-AGI solving sessions → Full logs of agent reasoning steps, program attempts, and evaluations
Tags
Tech Stack
Comments (0)
No comments yet
Editorially curated, with community endorsements as a secondary signal. Corrections welcome.