
[NEW RESEARCH] A better harness cuts AI cost and duration by 40%+, at the same quality, on any model. We benchmarked models from 6 leading labs across 22 tasks, comparing a harness optimized for token efficiency against a naive harness optimized only for completing the task. The efficient harness won on cost and speed by over 40%, with no quality tradeoff. Two findings stood out: → 𝐓𝐡𝐞 𝐞𝐟𝐟𝐢𝐜𝐢𝐞𝐧𝐜𝐲 𝐠𝐚𝐢𝐧 𝐢𝐬 𝐮𝐧𝐢𝐯𝐞𝐫𝐬𝐚𝐥. Every model got cheaper under the optimized harness, by 33% to 61%, regardless of its size or capability. → 𝐓𝐡𝐞 𝐪𝐮𝐚𝐥𝐢𝐭𝐲 𝐠𝐚𝐢𝐧 𝐢𝐬 𝐧𝐨𝐭. How much quality a model extracted from the same harness scaled almost perfectly with the model's baseline strength. Stronger models got proportionally more value out of harness optimization, not just better results. We call this pairing 𝐡𝐚𝐫𝐧𝐞𝐬𝐬 𝐥𝐞𝐯𝐞𝐫𝐚𝐠𝐞. This is the thinking behind Palmyra X6 and the platform updates we shipped alongside it, including the ability to choose the right model for the task, not just the biggest one. Read the full research and see what's new →

design, not model choice, cut cost and duration 33 to 61 percent at equal quality across six labs' models. The efficiency gain was universal; the quality gain scaled with the model's baseline strength.
Checking sign-in…
Loading comments…