math0_shared32k_d125m_ctx1024

d125m (134.1M parameters) on OpenWebMath (12 shards) (816M words, 0.86 passes), trained on RTX 3060 12GB (Home GPU). Started 28 Sep 2026, ended 28 Sep 2026.

Hypothesis

EN1-full's recipe and budget, trained on OpenWebMath (816M words of mathematical web pages), gives a math model that beats the English model on maths text.

What we learned

Yes, against the English model: 11% better on GSM8K, 32% on held-out OpenWebMath, 63% on competition maths (mostly LaTeX). Dropping every test item with copied text moves it under 2%: not memorisation.

Scores

Lower is better on the headline; change against the parent run, EN1-full.

OpenWebMath held-out, clean
1.06595
bits per byte · headline
Validation, in training
1.0307
bits per byte
−9.0%
Every published score of MATH0, against its parent EN1-full
Test setMATH0EN1-full (parent)Change
OpenWebMath held-out, clean headline
Mathematical web pages never used in training, minus every page with copied passages in the training text.
1.06595 —
GSM8K test
Grade-school maths word problems with worked solutions, scored as text.
1.07016 —
GSM8K exact match
The share of the 1,319 GSM8K test problems the model answers exactly right, with three worked examples in the prompt.
0.91% —
MATH test, clean
Competition problems with solutions, mostly in LaTeX, minus any with copied text in the training pages.
0.86027 —
OpenWebMath held-out, all
Every held-out OpenWebMath page, including the ones with passages copied into training pages.
0.99258 —
MATH test, all
All 5,000 MATH test problems, including the ones with copied text in the training pages.
0.84625 —

Bits per byte: how many bits the model needs, on average, to predict each byte of text it never saw in training. Lower is better, and it compares models with different tokenizers fairly, because every model is charged for the same bytes.

Exact match: the share of problems answered exactly right, as a percentage. Higher is better.

Curves

Training loss

Cross-entropy per token, by step. The first few percent of the run, far higher, run off the top; hover or the table has every value.

Held-out bits per byte

The validation split, evaluated during training, by step. Lower is better.

Throughput

Tokens per second, by step.

Model

Preset
d125m
Parameters
134,105,856
Outside embeddings
84,953,856
Layers · heads · width
12 · 12 · 768
Tokenizer
Shared unigram 32k

Muon + AdamW, rope, qk_norm, relu^2, untied head, context 1024

Data

Slice
OpenWebMath (12 shards)
Words
816,000,000
Training tokens
1,555,523,031
Passes
0.86 passes
Tokens seen
1,332,314,112

Compute

GPU
RTX 3060 12GB
Where
Home GPU
Steps
27,106 / 27,106
GPU hours
14.82
Cost
—
Spot restarts
—

Lineage

homed125m