
A thoughtful engagement with Sutton's RL-centric view of AI, arguing imitation learning and RL are complementary rather than opposed and unpacking the continual-learning bottleneck — useful framing for anyone reasoning about where learning is headed.
“Imitation learning is just short horizon RL. The episode is a token long.”
Dwarkesh Patel
“As planes are to birds, supervised learning might be to human cultural learning.”
Dwarkesh Patel
“It’s a bit like saying to someone pasteurizing milk, “Hey stop boiling that milk - we eventually want to serve it cold!””
Dwarkesh Patel
“Evolution does meta-RL to make an RL agent. That agent can selectively do imitation learning.”
Dwarkesh Patel
“If the LLMs do get to AGI first, the successor systems they build will almost certainly be based on Richard’s vision.”
Dwarkesh Patel
Checking sign-in…
Loading comments…