
Contamination-resistant benchmarks are the only credible way to compare coding models as training sets absorb public repos.
“they're able to directly run git log and then go through the commit hashes and cherrypick the ones out that contain the golden patches”
James Shi
“it will go ahead and implement the synchronous part, but it may drop the asynchronous part. We observed this in roughly two out of three cloud rollouts”
James Shi
“you're not going to be coming in there with a to-do list uh telling it to oh first do this and then do this”
James Shi
“we find that the average size of our solution is five times the lines of code”
James Shi
“stronger models have a great tendency to want to test their own work”
James Shi
Checking sign-in…
Loading comments…