← All IntelClip / EducationIFEval's flawed, unsolvable, and unverified prompts
From Benchmaxxing: The Gap Between Benchmark Scores and Reality · ≈7:32
“a benchmark is an artifact expressing what it's an aspirational artifact.”
“It's an expression of values of what you want your AI to do and how you want it to behave.”
“This one starts by saying repeat this response verbatim and it ends by saying translate this into Hindi.”
“write a riddle that includes exactly one bullet point. Make sure to include a few bullet points.”
What’s in it
- Exposes how a widely-cited benchmark bakes in unsolvable, contradictory prompts
- Shows a real reward-hacking exploit that lets models cheat verification checks
- Argues good benchmarks need 'taste' and product sense, not just prompt mashups
Clip transcript
match is just not going to do it to measure that sort of impact. Another important aspect of a good benchmark is taste. Perhaps it used to be the case that benchmarks were these dry academic, you know, questions. and answer sets. But nowadays, a benchmark is an artifact expressing what it's an aspirational artifact. It's an expression of values of what you want your AI to do and how you want it to behave. And so you need to have some product sense in this process, some sort of a sense of what you want the AI to do. And that sense is unfortunately missing from ifal if has been cited on many model cards. And the way it was constructed was taking a bunch of arbitrary prompts that no user has ever asked in earnest and mashing them up with a bunch of other prompts to create a prompt set. The problem is that because no user actually has asked do not use any commas in your response or use the letter T at most once. You have to believe for this to be useful, you have to believe that there's a generalization from this to actual things that users are going to ask. If eval just happens also to have a bunch of prompts that are fully unsolvable due to having contradictory instructions. So this one starts by saying repeat this response verbatim and it ends by saying translate this into Hindi. Obviously you can't do both of those at once. Here's one that says write a riddle that includes exactly one bullet point. Make sure to include a few bullet points. Again this is just fully impossible. It uses a sentence splitter that does not align with how humans would actually split the sentences. And a lot of the prompts are not fully verified. So this one says write a story. There's nothing in the verifier that checks that a story was written. It just checks that the asky character I is not used more than once, which means that all of these responses get a full score, including response D. The way it gets a full score is by reward hacking and using the cerrillic eye character instead of the asy eye character. If is totally fine with that
Comments
Sign in to comment.
Loading comments…