Stop Rationing Tokens: Let the Harness Pick the Model — Kimchi by Cast AI
Source
youtube.com
Author
AI Engineer
Date
Why it matters
It argues token rationing hurts developers and proposes measuring cost per task, with the harness routing between proprietary and open models. That is relevant to anyone managing coding-agent spend.
Key takeaways · AI-distilled
Cast AI built its own open-source coding agent harnessThe scaffolding around a model that turns it into a working agent — the loop, the tools it can call, and the rules for when to stop.Full definition →, Kimchi, after its coding-AI agentAn AI system that doesn't just answer once but works toward a goal in a loop — taking actions, reading the results, and deciding what to do next.Full definition → bill took off, and argues that limiting tokenThe chunk of text a model reads and writes in — roughly three-quarters of a word — and the unit AI usage is billed in.Full definition → cripples developers.
The harness chooses a proprietary or open model per task based on outcomes and switches automatically as new models ship; the talk presents three months of internal results from 300 employees.
Ferment runs milestone-based, self-scoring autonomous coding tasks for hours and deploys to staging. Teleport keeps agent sessions running in remote sandboxes after you close your laptop, and Studio adds a shared Kanban board for teams.
Terms in this piece · Glossary
agent harness — The scaffolding around a model that turns it into a working agent — the loop, the tools it can call, and the rules for when to stop.
AI agent — An AI system that doesn't just answer once but works toward a goal in a loop — taking actions, reading the results, and deciding what to do next.
token — The chunk of text a model reads and writes in — roughly three-quarters of a word — and the unit AI usage is billed in.