Vibeleaderboard
← All Intel
Intel / video

Yash Patel on the History and Future of AI Model Training

Source
Stanford Online
Author
Stanford Online
Date
Terms in this piece · Glossary
  • RLHF — Reinforcement learning from human feedback — training a model to prefer answers humans rate as better, which turns a raw text predictor into a usable assistant.
  • pretraining — The first, biggest phase of building a model: training it on enormous amounts of text so it learns language, facts, and reasoning in general.
Why it matters

Argues the next AI capability bottleneck is continual, data-efficient learning from sparse real-world feedback rather than more , and explains why verifiable domains like code and math currently dominate RL training — a framing worth understanding before betting on where model improvements come from next.

Key quotes

“These models today are not really like that.”

Yash Patel

“RL is kind of this like eval maxing machine.”

Yash Patel

“General models sort of set the floor, but in order to set the ceiling, you need to go and build, train models, create these specialized systems in order to differentiate yourself from all your competitors.”

Yash Patel

“My honest take is like scaling Transformers is working and as you know, there's a very simple recipe to be able to make these things smarter and better.”

Yash Patel
Read the source youtube.com
More from Stanford Online
Recommended reads
Comments

Checking sign-in…

Loading comments…