Math training runs

A 125M math model trained on mathematical web pages with the English recipe, and what fine-tuning on worked solutions adds.

The headline score is OpenWebMath held-out, clean. Mathematical web pages never used in training, minus every page with copied passages in the training text.

Bits per byte: how many bits the model needs, on average, to predict each byte of text it never saw in training. Lower is better, and it compares models with different tokenizers fairly, because every model is charged for the same bytes.

Exact match: the share of problems answered exactly right, as a percentage. Higher is better.

Every score on this track
OpenWebMath held-out, clean (lower is better)
Mathematical web pages never used in training, minus every page with copied passages in the training text.
GSM8K test (lower is better)
Grade-school maths word problems with worked solutions, scored as text.
GSM8K exact match (higher is better, a percentage)
The share of the 1,319 GSM8K test problems the model answers exactly right, with three worked examples in the prompt.
MATH test, clean (lower is better)
Competition problems with solutions, mostly in LaTeX, minus any with copied text in the training pages.
OpenWebMath held-out, all (lower is better)
Every held-out OpenWebMath page, including the ones with passages copied into training pages.
MATH test, all (lower is better)
All 5,000 MATH test problems, including the ones with copied text in the training pages.

Live now

Progress

Lower is better on the headline.

Best OpenWebMath held-out, clean over time

Each dot is a scored run, placed at the day it ended; the line is the best score so far. Numbered lines are this track's milestones, listed below.

Milestones
  1. 29 Sep 2026 · MATH0: first math model, 11% to 63% better than English on math
  2. 1 Oct 2026 · MATH0-SFT: fine-tuned on worked solutions, GSM8K 2.43%

Model size and OpenWebMath held-out, clean

Parameters (log scale) against OpenWebMath held-out, clean in bits per byte: this track's runs and other models measured the same way. Hover a point for its name.

Reference models

Other models measured on the same sets, for comparison. Notes say how they were measured.

Reference models scored on the Math track's test sets
ModelParametersOpenWebMath held-out, cleanGSM8K testGSM8K exact matchMATH test, clean
MuseMesh English 125M (EN1-full) Our English model, no maths training: same tokenizer, size, recipe and token budget as MATH0. The baseline MATH0 is compared with.134.1M1.570231.206651.59%2.30167

All Math runs

Published · 2 runs

All Math runs. Column headers sort the table.
Data
MATH0-SFT
math0_sft
1 Oct 2026 d125m
134.1M params
GSM8K + MATH worked solutions 0.22 — Keep — — 2.43%
MATH0
math0_shared32k_d125m_ctx1024
28 Sep 2026 d125m
134.1M params
OpenWebMath (12 shards)
816M words · 0.86×
14.82 — Winner 1.06595 1.07016 0.91%

Cost “—” means no cloud bill (our own GPU at home) or a cost that was not recorded; “Free” is free cloud compute. GPU-h is GPU hours.