Clip transcript
folks. Feedback and evals. So, here the quote is, "Shipping is the start, not the finish." So, what we do here uh on the agents team is we have kind of multiple ways we do evals and collect feedback. Um Obviously, you know, we'll have folks uh call in or or email us or just let us know and tell us, but the main way is we have this thumbs up and thumbs down mechanism and here someone is able to tell us, "Hey, this this worked really well. This was a great response." Or, "That wasn't a great response." And that signal we take and we're able to iterate on. Uh, and we can take it back and help improve, uh, you know, the response in the future. Um, we also have automated evals, so in in the in our CI we we have evals that run against real completions, so we could test the prompt against, "Hey, did it hit some tools? Did it do what it's supposed to do?" And that also helps with our accuracy. So, uh, those automated evals in conjunction with collecting feedback really help us, um, improve our uh, our our tools, our skills, um, our harness and and that's really how how we're able to iterate so fast and so quickly. Humans in the loop. So, this is a really cool feature we built where we deterministically interrupt the agent loop if there is a tool call approval required. So, if an agent tries to make a tool call that it needs human approval for, it'll show this UI and the human uh, can click accept or reject. So, explicitly rejecting or explicitly accepting, uh, the action that the agent is trying to make. And this ensures that, uh, you know, we're building trust and also ensuring that, uh, you know, we're being safe, especially when the agent's trying to do a mutating operation and always always always making sure that, um, humans are in the driver's seat.