Introducing Step Code v0.1.0.
Swift execution. High token efficiency. Long-horizon reliability.
Now open source under the MIT License.
Step Code handles the full development loop—from reading and editing code to running tests and shipping—from one CLI.
- 80.9% on Terminal-Bench 2.1 in evaluation
- 73.3% on Multi-Frame, our 150-task long-horizon benchmark
- One-command static site publishing with StepPage
GitHub:
benchmark — A standard public test set for comparing AI models — the shared scoreboards behind every "model X beats model Y" claim.
AI agent — An AI system that doesn't just answer once but works toward a goal in a loop — taking actions, reading the results, and deciding what to do next.
Why it matters
Step Code is a new open-source coding AI agentAn AI system that doesn't just answer once but works toward a goal in a loop — taking actions, reading the results, and deciding what to do next.Full definition → CLI with published long-horizon benchmarkA standard public test set for comparing AI models — the shared scoreboards behind every "model X beats model Y" claim.Full definition → results, giving practitioners another MIT-licensed option to evaluate against Claude Code and Codex.