EN1-muon

English DoneKeep

en1_muon_d125m_ctx1024

d125m (110.3M parameters) on FineWeb sample (4 files), trained on RTX 3060 12GB (Home GPU). Started 22 Sep 2026, ended 22 Sep 2026.

Hypothesis

Muon, an optimiser that orthogonalises each update, beats AdamW on the weight matrices at equal steps.

What we learned

Yes: FineWeb 1.1987 vs 1.2187 (-1.6%), pooled -1.4%. Web and wiki text gain; the character-level sets barely move. Muon went into the recipe.

Scores

Lower is better on the headline; change against the parent run, EN1-ctrl.

FineWeb validation
1.19873
bits per byte · headline
−1.6%
Validation, in training
1.2254
bits per byte
−1.5%
Every published score of EN1-muon, against its parent EN1-ctrl
Test setEN1-muonEN1-ctrl (parent)Change
FineWeb validation headline
About 15,000 FineWeb web documents never used in training, scored with a 1,024-token window.
1.19873 1.21874−1.6%
FineWeb validation, clean
The same documents minus any that overlap the training text.
1.21334 1.23271−1.6%
Pooled, five sets
All five English test sets pooled.
1.23271 1.25024−1.4%
WikiText-103 test
Good and featured Wikipedia articles, a standard language-modelling benchmark.
1.26125 1.28812−2.1%
enwik8 test
Raw Wikipedia markup; a character-level compression benchmark.
1.50046 1.5072−0.4%
text8 test
Lower-cased Wikipedia text with markup stripped; a character-level benchmark.
1.38396 1.37631+0.6%
FineWeb validation, during training
The validation number measured during training, the way the GPT-2 124M reference was measured, so the two can be compared.
1.2254 1.2441−1.5%

Bits per byte: how many bits the model needs, on average, to predict each byte of text it never saw in training. Lower is better, and it compares models with different tokenizers fairly, because every model is charged for the same bytes.

Curves

Training loss

Cross-entropy per token, by step. The first few percent of the run, far higher, run off the top; hover or the table has every value.

Held-out bits per byte

The validation split, evaluated during training, by step. Lower is better.

Throughput

Tokens per second, by step.

Model

Preset
d125m
Parameters
110,316,288
Outside embeddings
84,953,856
Layers · heads · width
12 · 12 · 768
Tokenizer
Shared unigram 32k

Muon + AdamW, learned positions, GELU, tied head, context 1024

Data

Slice
FineWeb sample (4 files)
Words
—
Training tokens
3,253,318,552
Passes
0.15 passes
Tokens seen
491,520,000

Compute

GPU
RTX 3060 12GB
Where
Home GPU
Steps
10,000 / 10,000
GPU hours
5.33
Cost
—
Spot restarts
—

Lineage

homed125m