Argues that building RL environments for each software agent skillA reusable instruction file that teaches an agent how to do one job well — the procedure, the tools, and what counts as done.Full definition → is inconsistent with near-term human-like continual learning. It helps practitioners judge how model capabilities will generalize.
Key takeaways · AI-distilled
Dwarkesh Patel argues labs' heavy spending on RL environments for skills like browsers and Excel reveals they expect models to keep generalizing poorly. A true humanlike learner would pick these up on the job, making the pre-baking pointless.
Patel's key crux: jobs are full of lab- or company-specific microtasks, such as spotting macrophages on one lab's slides. Building a custom training loop for each is not net productive; what is needed is AI that learns from semantic feedback and generalizes like a human.
Patel calls 'slow diffusion' cope: models truly like humans on a server would spread faster than hiring, with no lemons-market risk. He reads the gap between lab revenue and tens of trillions in knowledge-worker wages as a capability gap.
Patel warns against borrowing pretrainingThe first, biggest phase of building a model: training it on enormous amounts of text so it learns language, facts, and reasoning in general.Full definition →'s clean scaling trend to justify RL from verifiable reward, which has no well-known public trend. He cites an analysis of o-series benchmarks suggesting roughly a millionfold RL compute increase for a boost similar to one GPT generation.
Patel expects continual learning to arrive gradually, like in-context windowThe maximum amount of text a model can consider at once — its working memory for the current conversation or task.Full definition → learning after GPT-3: labs may ship something called continual learning soon, but human-level on-the-job learning may take 5 to 10 years, and rivals will likely replicate any early lead.
Terms in this piece · Glossary
agent skill — A reusable instruction file that teaches an agent how to do one job well — the procedure, the tools, and what counts as done.
pretraining — The first, biggest phase of building a model: training it on enormous amounts of text so it learns language, facts, and reasoning in general.
context window — The maximum amount of text a model can consider at once — its working memory for the current conversation or task.