Math training runs
A 125M math model trained on mathematical web pages with the English recipe, and what fine-tuning on worked solutions adds.
The headline score is OpenWebMath held-out, clean. Mathematical web pages never used in training, minus every page with copied passages in the training text.
Bits per byte: how many bits the model needs, on average, to predict each byte of text it never saw in training. Lower is better, and it compares models with different tokenizers fairly, because every model is charged for the same bytes.
Exact match: the share of problems answered exactly right, as a percentage. Higher is better.
Every score on this track
- OpenWebMath held-out, clean (lower is better)
- Mathematical web pages never used in training, minus every page with copied passages in the training text.
- GSM8K test (lower is better)
- Grade-school maths word problems with worked solutions, scored as text.
- GSM8K exact match (higher is better, a percentage)
- The share of the 1,319 GSM8K test problems the model answers exactly right, with three worked examples in the prompt.
- MATH test, clean (lower is better)
- Competition problems with solutions, mostly in LaTeX, minus any with copied text in the training pages.
- OpenWebMath held-out, all (lower is better)
- Every held-out OpenWebMath page, including the ones with passages copied into training pages.
- MATH test, all (lower is better)
- All 5,000 MATH test problems, including the ones with copied text in the training pages.
Live now
Progress
Lower is better on the headline.
Best OpenWebMath held-out, clean over time
Each dot is a scored run, placed at the day it ended; the line is the best score so far. Numbered lines are this track's milestones, listed below.
Milestones
Model size and OpenWebMath held-out, clean
Parameters (log scale) against OpenWebMath held-out, clean in bits per byte: this track's runs and other models measured the same way. Hover a point for its name.
Reference models
Other models measured on the same sets, for comparison. Notes say how they were measured.
| Model | Parameters | OpenWebMath held-out, clean | GSM8K test | GSM8K exact match | MATH test, clean |
|---|---|---|---|---|---|
| MuseMesh English 125M (EN1-full) Our English model, no maths training: same tokenizer, size, recipe and token budget as MATH0. The baseline MATH0 is compared with. | 134.1M | 1.57023 | 1.20665 | 1.59% | 2.30167 |
All Math runs
Published · 2 runs
| Data | |||||||||
|---|---|---|---|---|---|---|---|---|---|
| MATH0-SFT math0_sft |
1 Oct 2026 | d125m 134.1M params |
GSM8K + MATH worked solutions | 0.22 | — | Keep | — | — | 2.43% |
| MATH0 math0_shared32k_d125m_ctx1024 |
28 Sep 2026 | d125m 134.1M params |
OpenWebMath (12 shards) 816M words · 0.86× |
14.82 | — | Winner | 1.06595 | 1.07016 | 0.91% |
No runs match these filters. Clear the search or pick “All”.
Cost “—” means no cloud bill (our own GPU at home) or a cost that was not recorded; “Free” is free cloud compute. GPU-h is GPU hours.