English training runs

125M-class English models on FineWeb web text. The proving ground: optimisers and architecture are tested here first, against a known GPT-2 reference.

The headline score is FineWeb validation. About 15,000 FineWeb web documents never used in training, scored with a 1,024-token window.

Bits per byte: how many bits the model needs, on average, to predict each byte of text it never saw in training. Lower is better, and it compares models with different tokenizers fairly, because every model is charged for the same bytes.

Every score on this track
FineWeb validation (lower is better)
About 15,000 FineWeb web documents never used in training, scored with a 1,024-token window.
FineWeb validation, clean (lower is better)
The same documents minus any that overlap the training text.
Pooled, five sets (lower is better)
All five English test sets pooled.
WikiText-103 test (lower is better)
Good and featured Wikipedia articles, a standard language-modelling benchmark.
enwik8 test (lower is better)
Raw Wikipedia markup; a character-level compression benchmark.
text8 test (lower is better)
Lower-cased Wikipedia text with markup stripped; a character-level benchmark.
FineWeb validation, during training (lower is better)
The validation number measured during training, the way the GPT-2 124M reference was measured, so the two can be compared.

Live now

Progress

Lower is better on the headline.

Best FineWeb validation over time

Each dot is a scored run, placed at the day it ended; the line is the best score so far. Numbered lines are this track's milestones, listed below.

Milestones
  1. 22 Sep 2026 · EN0: the pipeline lands on the GPT-2 124M curve in English
  2. 23 Sep 2026 · EN1: Muon and a modern block stack, 4.8% better in English
  3. 24 Sep 2026 · EN1-full beats a plain AdamW GPT-2 124M at equal tokens

Model size and FineWeb validation, during training

Parameters (log scale) against FineWeb validation, during training in bits per byte: this track's runs and other models measured the same way. Hover a point for its name. The validation number measured during training, the way the GPT-2 124M reference was measured, so the two can be compared.

Reference models

Other models measured on the same sets, for comparison. Notes say how they were measured.

Reference models scored on the English track's test sets
ModelParametersFineWeb validation, during training
GPT-2 124M, AdamW (modded-nanogpt record) A published plain-AdamW GPT-2 124M run, read at the same 1.33B tokens on the same FineWeb validation documents; its own tokenizer, converted to bits per byte.124.4M1.1755

All English runs

Published · 7 runs

All English runs. Column headers sort the table.
Data
EN1-full
en1_full_d125m_ctx1024
23 Sep 2026 d125m
134.1M params
FineWeb sample (4 files)
0.41×
14.85 — Winner 1.10625 1.12346 1.13758
EN1-muon_arch_wd
en1_muon_arch_wd_d125m_ctx1024
23 Sep 2026 d125m
158.7M params
FineWeb sample (4 files)
0.15×
5.54 — Discard 1.1909 1.20451 1.2236
EN1-muon_arch
en1_muon_arch_d125m_ctx1024
22 Sep 2026 d125m
134.1M params
FineWeb sample (4 files)
0.15×
5.44 — Keep 1.16018 1.175 1.19241
EN1-arch
en1_arch_d125m_ctx1024
22 Sep 2026 d125m
134.1M params
FineWeb sample (4 files)
0.15×
5.16 — Keep 1.19028 1.20413 1.2237
EN1-ctrl
en1_adamw_ctrl_d125m_ctx1024
22 Sep 2026 d125m
110.3M params
FineWeb sample (4 files)
0.15×
4.95 — Keep 1.21874 1.23271 1.25024
EN1-muon
en1_muon_d125m_ctx1024
22 Sep 2026 d125m
110.3M params
FineWeb sample (4 files)
0.15×
5.33 — Keep 1.19873 1.21334 1.23271
EN0
en0_shared32k_d125m_ctx1024
21 Sep 2026 d125m
110.3M params
FineWeb sample (4 files)
0.41×
13.89 — Keep 1.15174 1.16785 1.18255

Cost “—” means no cloud bill (our own GPU at home) or a cost that was not recorded; “Free” is free cloud compute. GPU-h is GPU hours.