One Model Family, Two Gold-Level Results: Fine-Tuning Nemotron for IOI and IMO
- ID
- 32714
- Status
- summarized
- Published
- 07 Oct 2026, 8:45 PM
- Fetched
- 07 Oct 2026, 9:38 PM
- Provider
- Hugging Face Blog
- Category
- developer-ai
- Original URL
- https://huggingface.co/blog/nvidia/nemotron-ioi-and-imo-2026
- Source URL
- https://huggingface.co/blog/feed.xml
Summary
- Score
- 6.0
- Created
- 07 Oct 2026, 9:39 PM
- Tags
- Audience
- ai_ml_learnersdevelopersai_agent_users
What happened
NVIDIA authors on the Hugging Face blog describe fine-tuning its Nemotron 3 family into competition specialists: Nemotron-3-Ultra-CC with SFT plus a 'GenCorrect' loop scored 535.4/600 on IOI 2026, above the 361.12 gold threshold and the top human score of 498.27, while a generate-verify-refine system over Nemotron 3 Ultra SFT and RL checkpoints scored 30/42 on IMO 2026 against a gold threshold of 29. Training used 22,000 curated competitive-programming problems with synthetic reasoning traces, producing Nemotron-3-Nano-CC (30B total / 3B active, SFT + RL) and Nemotron-3-Ultra-CC (550B total / 55B active, SFT). The IOI run was an unofficial, unsupervised benchmark not included in the official ranking; the IMO proofs were graded by official IMO graders.
Why it matters
The transferable part is the four-step recipe, not the medals: curate domain problems and reasoning traces, apply SFT (and RL on the smaller Nano variant, 3B active parameters), then wrap the model in a generate-verify-refine inference loop instead of training a new foundation model. If you are scoping a domain specialist, this is evidence that a 3B-active model plus an inference loop is a plausible cheap path, and that the loop is doing real work. Treat the IOI 535.4 figure as a vendor-run, unofficial, unsupervised result — do not cite it as an official ranking. Nothing in the text is Malaysia- or SEA-specific.
Discussion angle
How much of that 535.4 came from the model versus the GenCorrect generate-verify-refine loop — and would the same loop lift a smaller open model you can actually run, given the Nano variant is 3B active parameters?