← All IntelClip / AI AgentsWorking with human raters: rubrics, examples, and explanations
From How Evals and Prompts Shape Agent Behavior — Preetika Bhateja & Daniel Bump, YouTube Ads · ≈7:28
“human human agreement should be strong within your team of what you consider a good use case and a good past case for an eval.”
AI Engineer
“when you do have your teams or other scale members rate eval if it's a pass or a fail, that doesn't really tell you much about where should the agent improve, what was the thinking that went behind coming to that conclusion.”
AI Engineer
“explanations really help you like get to the bottom of like where is it that the agent's actually like missing things. And then you can also use that input to train your agent better.”
AI Engineer
“providing them with a clear rubric of what they were actually rating with very clear examples”
AI Engineer
What’s in it
- Concrete practices for scaling human rating — clear rubric with worked examples up front, strong human-human agreement inside the team, and required written explanations so failures point at what to fix.
Clip transcript
raiders, LLM raiders, all of that. So we'll get a little bit more into what that looks like. So just a couple of things on like working with scale raiders and things that worked for us. Uh one was that providing them with a clear rubric of what they were actually rating with very clear examples. So we had a lot of situations, especially early on when you're building things. of course like there are so many edge cases and difficult cases that we've not tested out that a raider might encounter. So they're coming back to you saying oh what what what should I do in this case and then sometimes we as a team are like disagreeing on like should this be a pass should be should this be a fail things like that. So I think that's very important to do early on as much as clarity and examples you can give the raers that would be super helpful. So yeah to that point like human human agreement should be strong within your team of what you consider a good use case and a good past case for an eval. Uh the second things that we noticed that helped us a lot was getting explanations from raider. So when you do have your teams or other scale members rate eval if it's a pass or a fail, that doesn't really tell you much about where should the agent improve, what was the thinking that went behind coming to that conclusion. So it's helpful to get explanations of why they're rating something a certain way. And this is true for like if you do single side evals or sideby-side eval like when you're testing two models at the same time having explanations of why one thing failed or one thing worked can be super helpful. Uh other things to keep in mind is that you could also do like in our case it was multi output. So we were asking scale raiders um when we were building ads like are these ads accurate like did we do the right things for it? Is it brand safe? Is it like something that we expected it to be? Things of that nature. So we had like almost like a multi-turn eval system. If you're building those kind of cases, it can get a little tricky because it's not exactly a pass failure. Your raiders could be like, "Oh, well, it does very well in well in brand safety, but it does not do really good in like accuracy or things of that nature." So explanations really help you like get to the bottom of like where is it that the agent's actually like missing things. And then you can also use that input to train your agent better.
Comments
Sign in to comment.
Loading comments…