
Explains why agents that iterate against a validation set repeatedly don't degrade the way classical overfitting theory predicts, a finding that should change how much you trust an 's self-reported scores.
“Machine learning, at its core, is about generalization, not memorization.”
“Short descriptions cannot cheat because there isn't room.”
“The hill-climbing was extensive, but the thing that came out the other end was — or could have been — tiny.”
“What fits (into few tokens) doesn't overfit”
Checking sign-in…
Loading comments…