← All IntelClip / AI ToolsNeMo Tron's approach: pulling post-training SFT data into pre-training
From The Base Model Is Dead — Varun Singh, Arcee AI · ≈7:44
“it's really interesting to see the the NeMo Tron series leans so heavily into synthetic data, but uh MAI Thinking 1 kind of lean in the opposite direction.”
What’s in it
- Explains how NeMo Tron 3 Ultra blends post-training data into pretraining
- Shows a data recipe pie chart revealing SFT-labeled synthetic data upfront
- Contrasts NeMo Tron's synthetic-data-heavy approach with MAI Thinking 1's opposite strategy
Clip transcript
those. Um the other approach uh is to bring um post-training data and large-scale synthetic data back uh through pull it back through the process into the pre-training phase. Um the the bottom chart I've taken from NeMo Tron 3 Ultra. Um what they reveal there um the data recipe and not sure how readable it is, but these top three um on the left uh pie chart, the top three on the kind of right right side of it, uh they're all labeled SFT with with SFT as a prefix. And that's the type of question and answer kind of chat data set that you'd you'd expect to see only in post-training, but by pulling it back into the process, they're able to like get the model to learn um the shape of these conversations and what kind of tasks they might be expected to do downstream um from the very beginning of the pre-training process. Um and uh this follows like uh a similar um trend in like diminishing uh the amount of web text used in the model. Um Yeah. Um it's really interesting to see the the NeMo Tron series leans so heavily into synthetic data, but uh MAI Thinking 1 kind of lean in the opposite direction.
Comments
Sign in to comment.
Loading comments…