← All IntelClip / Education
A loss-value ladder mapping numbers to observable model behavior
From Training an LLM from Scratch, Locally — Angelos Perivolaropoulos, ElevenLabs · ≈49:51
Turns loss into a diagnostic: 4.17 is random, 3.3 means character frequencies, 1.5-2 produces words, 1.0-1.2 is coherent, below 1.0 signals overfitting on this dataset.
What’s in it
- Turns loss into a diagnostic: 4.17 is random, 3.3 means character frequencies, 1.5-2 produces words, 1.0-1.2 is coherent, below 1.0 signals overfitting on this dataset.
Clip transcript
at this point. Uh so, what we're we're going to start seeing is when we start from from big the beginning because it's a model we're training from scratch, the loss is going to be essentially random. And in that case, that would be natural log of 65, so it will start at around 4.17. Uh that basically means the model like knows nothing. It has no clue of what this data is. Uh and slowly we're going to start seeing the loss going down to 3.3. That's when the model is going to understand character frequencies. Uh it will still not be able to do words yet, but it might understand things like th as part of the, which is a common word. Like th is going to be part of the things that it starts generating. Then at around 2.5, it's going to it was going to get a little bit better about this th, and then it will understand the word in and stuff like that. Then at about 1.5 to two losses, it will start actually creating words. And then at around one 1.0 to 1.2, that's when the model is going to start being decent at at at this task. You will actually be able to understand names from the text. It will start creating things that start making sense. But then when the loss starts going below 1.0 for this specific data set, that's where we're going to start seeing overfitting. The model will still be like producing like reasonable things, but it will no longer start be getting better at it.
Recommended reads
Comments
Checking sign-in…
Loading comments…


