
Walks through how training loss scales predictably with model size, data, and compute, and how to allocate a fixed compute budget optimally between parameters and — essential for anyone planning or interpreting large training runs.
“The observation is simple in form: the training loss $L$ decreases predictably as we scale up model size $N$, dataset size $D$, and compute $C$, following a power-law curve, which appears as a straight line on a log-log plot.”
Checking sign-in…
Loading comments…