← All IntelClip / AI Agents
Verifiers beyond output checks: state, trace, and SME review
From From Agent Traces to Agent Simulations — Rustem Feyzkhanov, Snorkel AI · ≈10:21
Simulation verifiers should analyze final environment state, trace, and artifacts using deterministic checks plus LLM/agent judges, with subject-matter experts reserved for discrepancies.
What’s in it
- Simulation verifiers should analyze final environment state, trace, and artifacts using deterministic checks plus LLM/agent judges, with subject-matter experts reserved for discrepancies.
Clip transcript
simulate long-running horizon tasks. So, next part is verifiers. Like in traditional sense, usually when people speak about verifiers, you just verify the output. You get agent output, you verify it, that's it. That's why, let's say, how coding works. It's like we just verify the output code. We test, and so on. Here it's more complex. The way agent interacts with the simulation, we get a lot of different data. We get the world, basically how environment changed in the final environment state. What is your database state? What are the API responses? What were user replies? And so on. And your verifier analyzes final state, trace, and artifacts. So, how can you analyze it? Effectively, there are multiple ways to do it. You can have deterministic checks. Basically, uh and that can work really well for things like final output or tool calls, where like it's very easy to check whether it was correct or not. It Sometimes you can use LLM as a judge, or even harness as a judge, or agent as a judge to evaluate basically whether the uh trace quality was successful, whether um planning of the agent was correct, and so on. In this case, it really depends on the use case. So, you can use one or another or both depending on what's what works better. And finally, it's important to keep in mind that you can use sub- uh subject matter expert to review some of the traces, some of the outputs. Not for everything, but for cases where you see discrepancy in agent behavior and where you want human involvement.
Recommended reads
Comments
Checking sign-in…
Loading comments…

