
A concrete method for controlling spend at scale: build benchmarks from each agent's real work, re-pick models as the frontier shifts, and default subagents to a cheaper model while the primary plans and evaluates.
Checking sign-in…
Loading comments…