The study tested mini-SWE-agent and OpenCode with ten Qwen and DeepSeek open-weight models on SWE-benchThe standard benchmark for AI coding agents: real GitHub issues from real repositories, scored by whether the agent's patch passes the project's own tests.Full definition → Pro, ProgramBench and GitTaskBench, and found agent harnessThe scaffolding around a model that turns it into a working agent — the loop, the tools it can call, and the rules for when to stop.Full definition → value depends jointly on model capability and task type.
Complex harnesses gave diminishing gains on SWE-style issue repair as models got stronger, but could still help stronger models on more complex, open-ended repository-level tasks.
In component ablations on ProgramBench, structured tool useA model's ability to call external functions — run code, search the web, edit files — instead of only generating text.Full definition → and task-specific subagents gave the most stable gains, while context compactionSummarizing an agent's earlier conversation to free room in the context window so a long session can keep going.Full definition → and general-purpose subagents could hurt repository-generation performance.
Combining the helpful components in their NanoHarness lifted results over mini-SWE-agent by 7.37 points on Qwen3.7-Max and 6.21 points on DeepSeek-V4-Pro, recovering most of the gains of product-level harnesses, per the authors.
Terms in this piece · Glossary
SWE-bench — The standard benchmark for AI coding agents: real GitHub issues from real repositories, scored by whether the agent's patch passes the project's own tests.
agent harness — The scaffolding around a model that turns it into a working agent — the loop, the tools it can call, and the rules for when to stop.
context compaction — Summarizing an agent's earlier conversation to free room in the context window so a long session can keep going.
tool use — A model's ability to call external functions — run code, search the web, edit files — instead of only generating text.
Why it matters
Agent performance depends on the harness as well as the model. This study ablates tool registry, context compression, planning, subagents and lazy skills across ten open-weight models, showing which harness choices matter when building coding agents.