Yash Patel on the History and Future of AI Model Training
- Source
- Stanford Online
- Author
- Stanford Online
- Date
- RLHF — Reinforcement learning from human feedback — training a model to prefer answers humans rate as better, which turns a raw text predictor into a usable assistant.
- pretraining — The first, biggest phase of building a model: training it on enormous amounts of text so it learns language, facts, and reasoning in general.
Argues the next AI capability bottleneck is continual, data-efficient learning from sparse real-world feedback rather than more , and explains why verifiable domains like code and math currently dominate RL training — a framing worth understanding before betting on where model improvements come from next.
“These models today are not really like that.”
Yash Patel
“RL is kind of this like eval maxing machine.”
Yash Patel
“General models sort of set the floor, but in order to set the ceiling, you need to go and build, train models, create these specialized systems in order to differentiate yourself from all your competitors.”
Yash Patel
“My honest take is like scaling Transformers is working and as you know, there's a very simple recipe to be able to make these things smarter and better.”
Yash Patel
Checking sign-in…
Loading comments…





