Vibeleaderboard
← All Intel
Intel / article

Building Reliable Data Analytics Agents: Lessons from the KDD Cup

Source
developer.nvidia.com
Author
Jiwei Liu
Date
Why it matters

A small fixed LLM became a reliable data agent through harness design: one SQLite database, four tools, and a separate zero-temperature call for reading documents. Reusable for your own analytics agents.

Key takeaways · AI-distilled
  • NVIDIA's KGMON team placed second in the KDD Cup 2026 Data Agents competition, where every team had to use the same small, fixed . That made the around the model, not the model itself, the main thing to optimize.
  • KGMON loaded the task's CSV and JSON files into its existing SQLite database so the had one query surface, exposed through two custom functions, schema() and sql(query). The team says this cut routing failures and wasted turns.
  • Before the main loop, a read-only schema-scouting step briefed the agent on tables, likely join keys, look-alike fields, units, null patterns and row-grain issues, which the authors credit with fewer wrong-column and missed-join errors.
  • The agent could not read whole documents. It searched with capped previews or regex, then sent relevant chunks to prose_helper, a separate temperature-0 LLM call that returned a short answer or extracted a SQL table, keeping raw prose out of the main .
  • The team warns that autonomous improvement loops overfit through hardcoded training examples, contradictory prompt instructions and brittle postprocessing. Its guard: held-out tasks, prompt audits, trace review and human approval before a change is kept.
Terms in this piece · Glossary
  • AI agent — An AI system that doesn't just answer once but works toward a goal in a loop — taking actions, reading the results, and deciding what to do next.
  • LLM — A large language model — the neural network behind tools like Claude and ChatGPT, trained on huge amounts of text to predict what comes next.
  • agent harness — The scaffolding around a model that turns it into a working agent — the loop, the tools it can call, and the rules for when to stop.
  • context window — The maximum amount of text a model can consider at once — its working memory for the current conversation or task.
Read the source developer.nvidia.com
Recommended reads
Comments

Checking sign-in…

Loading comments…