
Most quality problems are problems misdiagnosed as model problems.
“Harness design alone can account for double-digit swings in benchmark results and significant differences in token cost, with the same underlying model.”
Ricardo Silveira Cabral and Paul Furgale
“NOOA takes a simpler approach: an agent is a single Python class . Its methods are its capabilities. Its fields are its state. Its docstrings are its prompts. Its type annotations are enforced contracts.”
Ricardo Silveira Cabral and Paul Furgale
“On SWE-bench Verified, NOOA reaches 82.2% with GPT-5.5 above the published leaderboard SOTA at submission (79.2%)—and 79.8% with Opus 4.6, using a general-purpose, 253-line agent with no benchmark-specific prompts.”
Ricardo Silveira Cabral and Paul Furgale
“With GPT-5.5, NOOA reaches 82.2% on SWE-bench Verified using 29 LLM calls and ~1.1M tokens per task. The comparison harnesses need 66 calls and 2.2M tokens to reach 78.2%, and 29 calls at 1.3M to reach 78.6%. Parity or better, at roughly half the cost.”
Ricardo Silveira Cabral and Paul Furgale
articleRun NVIDIA BioNeMo NIM Microservices for Protein Structure Prediction in Claude Science
articleSolving Agentic AI Fleet Challenges with NVIDIA Vera CPUChecking sign-in…
Loading comments…