ToolNavs Find Useful AI Tools
Submit Sign in
Back to AI information
Nemotron Takes Two Gold-Level Results: Past the Top Human at IOI, Over the Line at IMO

Nemotron Takes Two Gold-Level Results: Past the Top Human at IOI, Over the Line at IMO

AI information • Admin • • 5 views

Nemotron was summarized by the NVIDIA team in a Hugging Face blog post on October 7, 2026: two fine-tuned versions of the same model family reached gold-medal level at the 2026 International Olympiad in Informatics (IOI) and the International Mathematical Olympiad (IMO), with the IOI score also passing the top human contestant.

Two golds, judged separately

On the informatics side, the competitive-programming version answered live under the same time, no-internet, and submission limits as human contestants, scoring 535.4 out of 600, above the 361.12 gold threshold and the top human score of 498.27. The team is explicit, however, that this was an unofficial, unsupervised benchmark run and is not part of the official IOI ranking. On the mathematics side, the system wrote proofs in natural language only, with no formal prover, no external tools, and no internet access, scoring 30 out of 42 against an official gold threshold of 29, with full credit on four of six problems, and its submissions were graded by official IMO graders. The two results carry different kinds of weight and should be read separately.

Not a bigger base model

Both projects used the same playbook: start from Nemotron 3, apply supervised fine-tuning on domain data, add reinforcement learning where useful, and pair the specialist with an inference loop that generates, verifies, and refines candidate answers. The coding version learned from reasoning traces over 22,000 curated problems and used a strategy called GenCorrect to revise code across rounds of evaluator feedback. The math data covered writing proofs, refining them, verifying them, and judging verification itself, so the general model and the SFT and RL specialists could critique one another inside one system. The team's conclusion is that the medals came from combining fine-tuning with test-time compute, not from simply scaling parameters or sampling more attempts.

What is open, and the cost left unsaid

Released alongside are model checkpoints for both directions, training datasets, inference pipelines, and a new benchmark of 200 olympiad-level problems, hosted on Hugging Face and under NVIDIA's NeMo organization on GitHub. For researchers, the scores now come with a reproducible path rather than a screenshot. The other side deserves equal attention: the math system's final selection stage is deliberately compute-heavy, and a full rerun demands substantial compute and time. Open weights lower the barrier to participation, but they do not erase the compute gap, and that is worth remembering alongside the celebration of open competition results like these.

Recommended Tools

More