
Arbitrage
github.com/blader/arbitrage- Category
- AI Agents
- Rank
- No. 1922Tools index
Previous survey · No. 1930 ·
- Pricing
- Open Source
- Type
- TOOL
- Use case
- Coding
- Interfaces
- Agent Skill / Plugin
- Builder
- @blader
- GitHub
- 177 stars
- Date
About
Arbitrage is a Claude Code skill that optimizes AI token usage by routing expensive frontier-model tokens only to judgment tasks (planning, architecture, review) while dispatching all code-writing to the cheaper Codex agent running in the background. It includes a visual validation loop for frontend work and an escape hatch that falls back to the premium model after two failed Codex attempts. Think of it as cost-aware orchestration for agentic coding workflows.
What it does
A single instruction file that tells a Claude coding session to stop typing code itself. The session writes a spec with the exact test command that must pass, hands the writing to the OpenAI Codex command line agent, then inspects the returned diff and does all git work. For interface changes it runs the app, takes screenshots and sends concrete corrections back. There is no program here: the whole tool is a prompt plus a readme.
Why it's ranked here
The idea is sharp and cheap to try. It targets a real cost gap for people paying for both a Claude plan and a Codex plan, and the rules are specific: a routing table for which work stays and which goes, a two failed rounds limit before the main model takes over, and a list of common mistakes like dispatching without acceptance criteria. MIT licensed and installed with one clone. The weight is all in the wording, so nothing proves it saves money.
What's good
Every dispatch must name the test command that has to go green, which gives the reviewing model a clear pass or fail. Git stays with the main session, so the outside agent never commits or pushes. The skill refuses predicted failure: it tells the model to add pseudocode detail rather than assume the task is too hard. A red flags table answers the excuses a model tends to reach for, such as it being faster to write the code itself.
Tradeoffs
You need the Codex command line tool installed and a Codex subscription with spare quota; without both the skill has nothing to hand work to. It declares itself always active whenever coding, so it changes behavior on every task, including ones where writing a spec costs more than the edit. There are no tests, scripts or measurements in the repository, so any savings claim rests on the prose. It also assumes one isolated worktree per dispatch.
How to use it well
Fits someone who already runs Claude for planning and has an underused Codex plan, working in a repo with a real test suite so acceptance criteria mean something. Write specs as carefully as the skill asks, and split work at API boundaries rather than mid feature. It does not help with cost if you only pay for one vendor, and it adds no tooling for tracking what each round actually spent.
Technical notes+
The repository holds three files: README.md, SKILL.md and LICENSE. SKILL.md carries frontmatter that marks the skill always active in a Claude session and a routing table sending all implementation to codex via a /goal prompt, launched as codex exec --cd <worktree> with details kept in a spec file written into that worktree. Trivial one line edits, debugging analysis, visual validation and all git operations stay in the main session. The escape hatch triggers after a re-dispatch with corrective feedback still misses acceptance criteria. README.md gives clone targets for both the Claude and Codex skill directories.
Observed
- License
- MIT (LICENSE file in repository)
- Form
- Prompt-only skill: README.md, SKILL.md and LICENSE, no executable code
- Install
- git clone into the Claude Code or Codex skills directory
- Dependency
- Requires the codex CLI (codex exec) on PATH
- Activation
- SKILL.md frontmatter declares it always active when coding
- Tests
- No test directory or scripts in the repository tree
Read from README.md, SKILL.md, LICENSE.
What it can do
Route coding tasks to a cheaper AI model based on task type
A coding request or task description → Task dispatched to Codex for implementation or frontier model for judgment work
Generate code implementation using a background Codex agent
Code-writing task or feature specification → Completed code written by Codex at reduced token cost
Perform high-judgment tasks using the premium frontier model
Planning request, architecture decision, or diff review task → Plan, architecture spec, or reviewed diff produced by the premium model
Visually validate frontend UI output
Frontend code or UI component generated by Codex → Visual validation result confirming or rejecting the UI rendering
Automatically fall back to the premium model after repeated Codex failures
Two consecutive failed Codex implementation attempts → Task re-executed by the premium frontier model as an escape hatch
Orchestrate multi-agent agentic coding workflows
A complex coding workflow with multiple subtasks → Coordinated task execution split across Codex and frontier model based on cost and complexity
Reduce AI token costs for coding projects
A software development task → Completed task using minimized frontier-model token consumption
Intel on Arbitrage
Tags
Comments (0)
No comments yet
Editorially curated, with community endorsements as a secondary signal. Corrections welcome.