Two Gold-Level Results, One Model Family: Nemotron At IOI And IMO
AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: Two Gold-Level Results, One Model Family: Nemotron At IOI And IMO on ThorstenMeyerAI.com

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get tech for your team delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

Hugging Face says two specialized systems from its Nemotron 3 family scored 535.4 out of 600 at the 2026 International Olympiad in Informatics and 30 out of 42 at the International Mathematical Olympiad. IMO graders officially awarded the proof score; the IOI result came from an unofficial run and did not count in the competition rankings.

Hugging Face says two specialized systems built from its Nemotron 3 model family reached gold-level scores at the 2026 International Olympiad in Informatics (IOI) and International Mathematical Olympiad (IMO), as detailed in the original analysis. The company reported 535.4 points out of 600 at IOI and 30 out of 42 at IMO; the IMO proofs were graded by official graders, while the IOI score came from an unofficial run and was excluded from the official ranking.

For the programming contest, Hugging Face used a competition-specific version of Nemotron-3-Ultra-CC, trained with supervised fine-tuning (SFT) and paired with GenCorrect. The company describes GenCorrect as an iterative method that generates candidate solutions, evaluates them and refines them. Hugging Face says the system ran prospectively under the same time, internet-access and submission constraints as human contestants. Its reported score exceeded the 361.12-point gold threshold and the top human score of 498.27, but the run was unsupervised and unofficial.

For the IMO, the team combined the general Nemotron 3 Ultra model with SFT and reinforcement-learning checkpoints in a system that wrote proofs in natural language. It generated candidate proofs, scored and critiqued them, then revised promising attempts. According to Hugging Face, official graders awarded the submissions 30 points, above the stated 29-point gold threshold, including full credit on four of the six problems. The company says this system used no formal prover, external tools or internet access.

The projects used different training material and specialist methods. Hugging Face reports that the IOI work drew on 22,000 programming problems and synthetic reasoning traces. The IMO SFT data contained 414,890 quality-filtered examples from 15,818 proof problems; its reinforcement-learning model was trained on 9,597 problems selected near the model’s capability frontier. These figures, like the scores, come from the company’s account.

At a glance
reportWhen: Reported after the 2026 competitions; t…
The developmentHugging Face reported that specialized Nemotron systems achieved scores above the stated gold thresholds at the 2026 IOI and IMO, with different levels of official validation.
At a glance
reportWhen: Reported after the 2026 competitions
The developmentHugging Face reported that systems fine-tuned from Nemotron 3 scored above the gold thresholds at IOI 2026 and IMO 2026.

Two Subjects, Unequal Validation

Hugging Face reported results from adapting systems in its Nemotron model family for competitive programming and mathematical proof writing. The company says its methods included task-specific fine-tuning and additional inference-time steps to assess and revise candidate answers. It also says SFT and reinforcement-learning checkpoints contributed to the IMO system, while GenCorrect was used to refine IOI solutions.

The scores represent different kinds of evidence. The IMO proofs received official grading; the IOI result was a company-reported benchmark run and did not constitute an official placement or medal. The scores do not by themselves establish performance across broader programming workloads, mathematical research or unfamiliar contest formats.

The supplied account does not isolate the contribution of each training and inference component. Independent evaluation would be needed to assess the IOI run and determine how well either system performs on other tasks.

Amazon

AI coding assistant for programming competitions

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

From IOI Experiments to Proofs

Hugging Face presents the 2026 work as an extension of its IOI 2025 experiments, which tested how post-training and additional computation at inference time affected coding scores. The company reported that a Nemotron-3-Nano-CC model rose from 130 points before post-training to 280 after SFT and 291 after reinforcement learning. With GenCorrect, it reached 468, above that year’s stated gold threshold of 438.3. An Ultra-CC version scored 502 using the same test-time strategy.

The 2026 efforts applied related ideas to two different forms of reasoning. IOI problems require programs that work against executable tests, including hidden tests; IMO problems require written arguments that meet standards of mathematical rigor. Hugging Face says its IMO system drew on both SFT and reinforcement-learning checkpoints rather than relying on a single model version. That makes the two results related through their model family and development approach, but not directly comparable as one combined benchmark.

“Success at both points to something broader.”

— Hugging Face

Amazon

Mathematical proof writing AI tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Independent Checks Still Needed

The IOI score was not part of the official ranking, and the supplied account does not describe an independent audit or replication of that run. Although Hugging Face says the system operated prospectively under contest-like constraints, readers do not have enough information here to assess how the run was monitored or independently verified.

The results also come from the team that built the systems. The source material does not establish how well performance transfers to other competitions, unseen proof styles or real-world coding tasks, nor does it fully separate the effects of training data, fine-tuning, reinforcement learning and inference-time search. The company reports releases for some IMO materials, but the available account does not provide a timetable for further releases or independent evaluation of the IOI system.

Amazon

AI system for competitive programming

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Checkpoints and Reproduction

Hugging Face says its Nemotron Labs IMO 2026 collection includes SFT and reinforcement-learning checkpoints, both training datasets and Nemotron-IMO-Bench, a benchmark with 200 olympiad-level problems. The company also points to an IMO paper describing the training and generate-verify-refine system, as well as a NeMo-Skills repository.

Those materials may allow researchers to inspect the approach and run tests beyond the reported contest problems. Further scrutiny would be needed to establish whether outside teams can reproduce the scores, particularly the unofficial IOI result. The source gives no schedule for an independent review or additional evaluations.

Amazon

Natural language proof generator

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What scores did the Nemotron systems report?

Hugging Face reported 535.4 out of 600 at IOI and 30 out of 42 at IMO. The company said both exceeded their stated gold thresholds, but only the IMO proofs received official grading.

Did Nemotron win an official IOI medal?

No official IOI medal or ranking is reported. The 535.4-point result came from an unofficial run and was not entered in the competition’s official ranking.

How was the IMO result evaluated?

Hugging Face says official IMO graders assessed the submitted proofs and awarded 30 of 42 points, with full credit on four of six problems. The company states that the system did not use a formal prover, external tools or internet access.

What methods did the systems use?

The IOI system paired supervised fine-tuning with GenCorrect, which generates, evaluates and refines candidate programs. The IMO system combined a general Nemotron 3 Ultra model with SFT and reinforcement-learning checkpoints to generate and revise written proofs.

Can researchers test the IMO system?

Hugging Face says it released an IMO collection with model checkpoints, training datasets and Nemotron-IMO-Bench, a 200-problem benchmark. The supplied source does not give details on independent replication of either competition result.

Primary source: Hugging Face · via ThorstenMeyerAI.com

HALLOWEEN

Halloween Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The Local-First Agentic Operator

A single operator, leveraging agentic AI, now builds and manages multiple complex products, previously requiring entire organizations, marking a shift in software development.

The AI Restriction Nobody Saw In China’s Optical-Transceiver Revolution

A draft US measure targets Chinese optical transceivers, highlighting the strategic importance of data interconnects in AI infrastructure and revealing unresolved issues.

IdeaClyst: The Engine That Decides What’s Worth Building

IdeaClyst is an idea engine that helps founders identify valuable product opportunities by analyzing roadmaps and market data, transforming rough concepts into actionable plans.

AI Workloads Threaten to Drain America’s Electrical Capacity

Dramatic increases in AI workloads threaten to drain America’s electrical capacity, raising concerns about future power shortages and infrastructure resilience.