Clip transcript
and how it should have performed. Once we collect a a a relatively good number of traces, once we have um enough information, we basically run an agent workflow that we built. So, basically, we analyze all the traces with both the negative and positive feedback, but we are more interested about the negative feedback here. So, we analyze all these traces. We do clustering of the failure modes, right? So, based on all the feedback that we collected, we try to understand, okay, given all this five feedback, what are like the five, six, seven failure modes that we have here? And then, we analyze these um we analyze these clusters of failure modes with our subject matter experts. So, we say, "Okay, based on those 10 traces, we can note that the agent in this specific case performed poorly because of reason X, for example." We validate all this with our subject matter experts. And then we do have a fix proposal, which gets generated and implemented by our coding agent. Then we implement this and then we ship it to production to see whether it actually improved on real data. And of course, we use the traces that we do have as regressions. So, we test our fix against our traces first of all. But let's see that in detail.