One Model Family, Two Gold-Level Results: Fine-Tuning Nemotron for IOI and IMO
Source
huggingface.co
Date
Why it matters
Shows how much of a gold-level result came from SFT, RL and an iterative generate-evaluate-refine loop on top of a general model. Teams building reasoning or coding agents can reuse the loop pattern.
Terms in this piece · Glossary
test-time compute — Spending more computation when the model answers — thinking longer, trying multiple attempts — to buy accuracy without training a bigger model.