Measuring the Microtask Eligibility Gap: When Is an Off-the-Shelf SLM Enough for an Agent Harness?
Source
Jundong Hu, Shekar Ramachandran
Author
Jundong Hu, Shekar Ramachandran
Date
Key takeaways · AI-distilled
The four microtasks were auto-approving shell commands, writing memory, selecting tools and ranking past turns. A configuration passed only if its confidence bound cleared a threshold anchored to a cheap non-LLMA large language model — the neural network behind tools like Claude and ChatGPT, trained on huge amounts of text to predict what comes next.Full definition → baseline.
4-bit quantizationShrinking a model by storing its numbers less precisely — like rounding — so it runs faster and fits on smaller hardware, at a small quality cost.Full definition → (RTN, GPTQ, AWQ) did damage that depended on model size and moved no configuration into eligibility, so the gap tracked model size more than numeric precision.
The result replicated on Llama-3.x (12 of 12 configurations ineligible) and held across reworded prompts, with 0 of 112 eligible over the original plus three neutral paraphrases per cell.
The authors recommend placing SLMs behind a baseline that meets the threshold and using them only where it falls short; a 4B re-ranker over a BM25 shortlist beat BM25 alone, though it did not itself certify as eligible.
Terms in this piece · Glossary
AI agent — An AI system that doesn't just answer once but works toward a goal in a loop — taking actions, reading the results, and deciding what to do next.
agent harness — The scaffolding around a model that turns it into a working agent — the loop, the tools it can call, and the rules for when to stop.
LLM — A large language model — the neural network behind tools like Claude and ChatGPT, trained on huge amounts of text to predict what comes next.
quantization — Shrinking a model by storing its numbers less precisely — like rounding — so it runs faster and fits on smaller hardware, at a small quality cost.
Why it matters
Before delegating agent harnessThe scaffolding around a model that turns it into a working agent — the loop, the tools it can call, and the rules for when to stop.Full definition → microtasks like shell auto-approval to small local models, this shows none of the tested Qwen3 sizes met cheap-baseline thresholds, and which failures are fixable by thresholding.