Vibeleaderboard
Index / tool
Visit github.com
Category
AI Agents
Rank
No. 1922Tools index

Previous survey · No. 1930 ·

Pricing
Open Source
Type
TOOL
Use case
Coding
Interfaces
Agent Skill / Plugin
Builder
@blader
GitHub
177 stars
Date

About

Arbitrage is a Claude Code skill that optimizes AI token usage by routing expensive frontier-model tokens only to judgment tasks (planning, architecture, review) while dispatching all code-writing to the cheaper Codex agent running in the background. It includes a visual validation loop for frontend work and an escape hatch that falls back to the premium model after two failed Codex attempts. Think of it as cost-aware orchestration for agentic coding workflows.

What it does

A single instruction file that tells a Claude coding session to stop typing code itself. The session writes a spec with the exact test command that must pass, hands the writing to the OpenAI Codex command line agent, then inspects the returned diff and does all git work. For interface changes it runs the app, takes screenshots and sends concrete corrections back. There is no program here: the whole tool is a prompt plus a readme.

Why it's ranked here

The idea is sharp and cheap to try. It targets a real cost gap for people paying for both a Claude plan and a Codex plan, and the rules are specific: a routing table for which work stays and which goes, a two failed rounds limit before the main model takes over, and a list of common mistakes like dispatching without acceptance criteria. MIT licensed and installed with one clone. The weight is all in the wording, so nothing proves it saves money.

What's good

Every dispatch must name the test command that has to go green, which gives the reviewing model a clear pass or fail. Git stays with the main session, so the outside agent never commits or pushes. The skill refuses predicted failure: it tells the model to add pseudocode detail rather than assume the task is too hard. A red flags table answers the excuses a model tends to reach for, such as it being faster to write the code itself.

Tradeoffs

You need the Codex command line tool installed and a Codex subscription with spare quota; without both the skill has nothing to hand work to. It declares itself always active whenever coding, so it changes behavior on every task, including ones where writing a spec costs more than the edit. There are no tests, scripts or measurements in the repository, so any savings claim rests on the prose. It also assumes one isolated worktree per dispatch.

How to use it well

Fits someone who already runs Claude for planning and has an underused Codex plan, working in a repo with a real test suite so acceptance criteria mean something. Write specs as carefully as the skill asks, and split work at API boundaries rather than mid feature. It does not help with cost if you only pay for one vendor, and it adds no tooling for tracking what each round actually spent.

Technical notes+

The repository holds three files: README.md, SKILL.md and LICENSE. SKILL.md carries frontmatter that marks the skill always active in a Claude session and a routing table sending all implementation to codex via a /goal prompt, launched as codex exec --cd <worktree> with details kept in a spec file written into that worktree. Trivial one line edits, debugging analysis, visual validation and all git operations stay in the main session. The escape hatch triggers after a re-dispatch with corrective feedback still misses acceptance criteria. README.md gives clone targets for both the Claude and Codex skill directories.

Observed

License
MIT (LICENSE file in repository)
Form
Prompt-only skill: README.md, SKILL.md and LICENSE, no executable code
Install
git clone into the Claude Code or Codex skills directory
Dependency
Requires the codex CLI (codex exec) on PATH
Activation
SKILL.md frontmatter declares it always active when coding
Tests
No test directory or scripts in the repository tree

Read from README.md, SKILL.md, LICENSE.

What it can do

  • Route coding tasks to a cheaper AI model based on task type

    A coding request or task description → Task dispatched to Codex for implementation or frontier model for judgment work

  • Generate code implementation using a background Codex agent

    Code-writing task or feature specification → Completed code written by Codex at reduced token cost

  • Perform high-judgment tasks using the premium frontier model

    Planning request, architecture decision, or diff review task → Plan, architecture spec, or reviewed diff produced by the premium model

  • Visually validate frontend UI output

    Frontend code or UI component generated by Codex → Visual validation result confirming or rejecting the UI rendering

  • Automatically fall back to the premium model after repeated Codex failures

    Two consecutive failed Codex implementation attempts → Task re-executed by the premium frontier model as an escape hatch

  • Orchestrate multi-agent agentic coding workflows

    A complex coding workflow with multiple subtasks → Coordinated task execution split across Codex and frontier model based on cost and complexity

  • Reduce AI token costs for coding projects

    A software development task → Completed task using minimized frontier-model token consumption

Intel on Arbitrage

More in Intel

Tags

claudecodextoken-optimizationagentic-codingllmcost-efficiencyorchestrationdeveloper-tools

Comments (0)

No comments yet

Editorially curated, with community endorsements as a secondary signal. Corrections welcome.