← All IntelClip / Developer ToolsDeepSWE v1.1 hardens against reward hacking
From DeepSWE: A Contamination-Resistant Coding Benchmark — James Shi, Datacurve · ≈15:18
“So in here we've taken some additional measures to guard against cheating uh reward hacking uh by ensuring you know the verifier runtime is fully separate now from the agent runtime.”
What’s in it
- Details anti-cheating fixes in DeepSuite v1.1 eval environment
- Explains how verifier and agent runtimes were separated
- Covers new safeguards against reward hacking in coding benchmarks
Clip transcript
models performance would also be a great addition here. So we've already uh released deep suite v 1.1. So in here we've taken some additional measures to guard against cheating uh reward hacking uh by ensuring you know the verifier runtime is fully separate now from the agent runtime. Um also making sure the test reports are in a more standardized format and also making sure that we've trimmed all of the git refs and the commits besides the uh base commit that our agents are working on. So all of this in service of just making the environments more robust and more
Comments
Sign in to comment.
Loading comments…