benchmark — A standard public test set for comparing AI models — the shared scoreboards behind every "model X beats model Y" claim.
Why it matters
Shows a fine-tuned small model beating scaffolded frontier models on text-to-SQL, and names the reward bugs that quietly corrupt RLVR runs: execution match mistaken for semantic equivalence, and knowledge-blind outcome rewards.