benchmark — A standard public test set for comparing AI models — the shared scoreboards behind every "model X beats model Y" claim.
fine-tuning — Taking a trained model and training it a bit more on your own examples so it gets better at one specific job.
Why it matters
RL fine-tuningTaking a trained model and training it a bit more on your own examples so it gets better at one specific job.Full definition → lifted gpt-oss-120b from 38.6 to 45.8 on the Metaculus AI benchmarkA standard public test set for comparing AI models — the shared scoreboards behind every "model X beats model Y" claim.Full definition →, matching frontier models while staying decorrelated from them, which is what makes it valuable inside an ensemble.