← All IntelClip / EducationReward hacking and adversarial verifier design
From Benchmaxxing: The Gap Between Benchmark Scores and Reality · ≈5:55
“Reward hacking is basically when a model finds a lazy and creative way to meet the letter of the law, but not the spirit.”
“Gradient descent is basically like water flowing downhill looking for the path of least resistance.”
“You need to think about designing your rewards as a adversarial process against this maximally lazy agent.”
What’s in it
- Explains reward hacking: models gaming rules instead of intent
- Frames reward design as an adversarial game against lazy agents
- Uses a water-flowing-downhill analogy for gradient descent optimization
Clip transcript
information. Reward hacking is also a big problem. Reward hacking is basically when a model finds a lazy and creative way to meet the letter of the law, but not the spirit. You need to think about designing your rewards as a adversarial process against this maximally lazy agent. Gradient descent is basically like water flowing downhill looking for the path of least resistance. And so your verifiers need to be robust to that.
Comments
Sign in to comment.
Loading comments…