Vibeleaderboard
← All Intel
Intel / article

Tracing Agent Harness Behavior with NVIDIA NeMo Relay

Source
William Markito Oliveira
Author
William Markito Oliveira
Date
Key takeaways · AI-distilled
  • An ATIF trajectory shows which tool the model asked to run, not whether it worked. The authors say to confirm outcomes in the ATOF stream, where matching start and end events share a uuid and parent_uuid links a tool call to its parent.
  • NVIDIA's recipe for judging a change: alter one thing, hold model snapshot, provider, input, budget and timeout fixed, run equal repetitions per arm, and compare verified task outcomes before looking at calls, or cost.
  • In the 108-run ToolPerf rerun, Claude Sonnet 4.5 was effectively unchanged (24/27 vs 23/27). Qwen3 Coder 30B rose from 19/27 to 22/27, but mean calls went from 3.8 to 4.9, tool data from 16 KB to 33 KB and duration from 27 s to 42 s.
  • The task audit shows the trade-off: recovery recipes lifted Qwen on a blocked-command task from 33% to 100%, while zero-match probe output on a case-insensitive search pushed turns from 3.3 to 9.3, a regression the authors flag.
  • The authors caution that runs using live web search are for exploring behavior, not ranking models; a fair comparison needs fixed search responses and repeated runs. Traces can contain prompts, tool arguments and file paths, so review them before sharing.
Terms in this piece · Glossary
  • AI agent — An AI system that doesn't just answer once but works toward a goal in a loop — taking actions, reading the results, and deciding what to do next.
  • agent harness — The scaffolding around a model that turns it into a working agent — the loop, the tools it can call, and the rules for when to stop.
  • token — The chunk of text a model reads and writes in — roughly three-quarters of a word — and the unit AI usage is billed in.
  • LLM — A large language model — the neural network behind tools like Claude and ChatGPT, trained on huge amounts of text to predict what comes next.
Why it matters

Shows how to use ordered traces to test whether a harness change actually improves outcomes. In the case study a model recovered more tasks but used more calls, data and latency.

Read the source developer.nvidia.com
Recommended reads
Comments

Checking sign-in…

Loading comments…