
A deep technical conversation with two Anthropic researchers on how far RL post-training can actually scale, where continual learning and compute bottleneck AGI, and how mechanistic interpretability traces a model's reasoning — useful for engineers reasoning about capabilities and near-term autonomy.
“Okay, so I think the biggest thing that's changed is that RL in language models has finally worked. We finally have proof of an algorithm that can give us expert human reliability and performance, given the right feedback loop.”
Sholto Douglas
“I really do think by the end of this year to this time next year, we will have software engineering agents that can do close to a day's worth of work for a junior engineer”
Sholto Douglas
“I actually think a Nobel Prize is more likely than a Pulitzer Prize-winning novel in some respects.”
Sholto Douglas
“You're going to spend literally weeks giving them feedback whereas we'll give up on a model in minutes.”
Dwarkesh Patel
“If you look at the circuit, you can see that it's not actually doing any of the math, it's paying attention to that you think the answer's four”
Trenton Bricken
Checking sign-in…
Loading comments…