
A rare first-party account of the failure modes inside a large run, including the bugs that silently flatten a loss curve and the replica-equality check that catches them.
“the rule of thumb is if task is too hard for your model, then your model will start to fall on its face. Lose correctness, lose diversity.”
Marah Abdin
“synthetic data gives us a track to extract some of these features and project them on some new planes”
Marah Abdin
“If you've got data that sucks, you can't train a good model. If you've got a training code base that sucks, you also can't.”
Robert McHardy
“we noticed about 0.5% of the gradient gets silently corrupted, essentially replaced by random values”
Robert McHardy
“The recipe held, it scaled, and we will continue scaling it from here.”
Robert McHardy
Checking sign-in…
Loading comments…