
If you're building RL environments or agents on tasks without clean ground truth, this offers a concrete alternative — using judge models, QA from real documents/repos, and a reverse hide-and-reveal trick for generating difficulty-tunable tasks — plus a clear warning about reward hacking pitfalls to watch for.
“environments and evals are really the same thing”
Will Brown
“reward hacking is when you have a kind of loose proxy for your objective that is undefined at the boundaries”
Will Brown
“You can verify the easy problem and then learn on the hard problem.”
Will Brown
“RL's great for refining skills, but less so for incorporating like dense new knowledge.”
Will Brown
“classical machine learning will tell you you can train for the distribution, but generalizing outside of the distribution is kind of an undefined problem”
Will Brown
video"My name is... my name is...": A Linguistic Map for Voice Agents — Midam Kim, ServiceNow
videoAct, Confirm, or Stop? Smarter behavior for AI assistants, wearables & robots — Amit Desai, Roku
videoSpeech-to-Speech Model Research at Google DeepMind — Valeria Wu Fon & Tom Ouyang, Google DeepMind
videoVoice Agents Can Just Do Things — Charlie Guo, OpenAIChecking sign-in…
Loading comments…