Why RL Won — Kyle Corbitt, OpenPipe (acq. CoreWeave)
Source
youtube.com
Author
Latent Space
Date
Why it matters
A practitioner who built an RL and fine-tuningTaking a trained model and training it a bit more on your own examples so it gets better at one specific job.Full definition → product explains why RL became the method that stuck, which informs when to fine-tune agents.
Terms in this piece · Glossary
fine-tuning — Taking a trained model and training it a bit more on your own examples so it gets better at one specific job.
AI agent — An AI system that doesn't just answer once but works toward a goal in a loop — taking actions, reading the results, and deciding what to do next.