English training runs
125M-class English models on FineWeb web text. The proving ground: optimisers and architecture are tested here first, against a known GPT-2 reference.
The headline score is FineWeb validation. About 15,000 FineWeb web documents never used in training, scored with a 1,024-token window.
Bits per byte: how many bits the model needs, on average, to predict each byte of text it never saw in training. Lower is better, and it compares models with different tokenizers fairly, because every model is charged for the same bytes.
Every score on this track
- FineWeb validation (lower is better)
- About 15,000 FineWeb web documents never used in training, scored with a 1,024-token window.
- FineWeb validation, clean (lower is better)
- The same documents minus any that overlap the training text.
- Pooled, five sets (lower is better)
- All five English test sets pooled.
- WikiText-103 test (lower is better)
- Good and featured Wikipedia articles, a standard language-modelling benchmark.
- enwik8 test (lower is better)
- Raw Wikipedia markup; a character-level compression benchmark.
- text8 test (lower is better)
- Lower-cased Wikipedia text with markup stripped; a character-level benchmark.
- FineWeb validation, during training (lower is better)
- The validation number measured during training, the way the GPT-2 124M reference was measured, so the two can be compared.
Live now
Progress
Lower is better on the headline.
Best FineWeb validation over time
Each dot is a scored run, placed at the day it ended; the line is the best score so far. Numbered lines are this track's milestones, listed below.
Milestones
Model size and FineWeb validation, during training
Parameters (log scale) against FineWeb validation, during training in bits per byte: this track's runs and other models measured the same way. Hover a point for its name. The validation number measured during training, the way the GPT-2 124M reference was measured, so the two can be compared.
Reference models
Other models measured on the same sets, for comparison. Notes say how they were measured.
| Model | Parameters | FineWeb validation, during training |
|---|---|---|
| GPT-2 124M, AdamW (modded-nanogpt record) A published plain-AdamW GPT-2 124M run, read at the same 1.33B tokens on the same FineWeb validation documents; its own tokenizer, converted to bits per byte. | 124.4M | 1.1755 |
All English runs
Published · 7 runs
| Data | |||||||||
|---|---|---|---|---|---|---|---|---|---|
| EN1-full en1_full_d125m_ctx1024 |
23 Sep 2026 | d125m 134.1M params |
FineWeb sample (4 files) 0.41× |
14.85 | — | Winner | 1.10625 | 1.12346 | 1.13758 |
| EN1-muon_arch_wd en1_muon_arch_wd_d125m_ctx1024 |
23 Sep 2026 | d125m 158.7M params |
FineWeb sample (4 files) 0.15× |
5.54 | — | Discard | 1.1909 | 1.20451 | 1.2236 |
| EN1-muon_arch en1_muon_arch_d125m_ctx1024 |
22 Sep 2026 | d125m 134.1M params |
FineWeb sample (4 files) 0.15× |
5.44 | — | Keep | 1.16018 | 1.175 | 1.19241 |
| EN1-arch en1_arch_d125m_ctx1024 |
22 Sep 2026 | d125m 134.1M params |
FineWeb sample (4 files) 0.15× |
5.16 | — | Keep | 1.19028 | 1.20413 | 1.2237 |
| EN1-ctrl en1_adamw_ctrl_d125m_ctx1024 |
22 Sep 2026 | d125m 110.3M params |
FineWeb sample (4 files) 0.15× |
4.95 | — | Keep | 1.21874 | 1.23271 | 1.25024 |
| EN1-muon en1_muon_d125m_ctx1024 |
22 Sep 2026 | d125m 110.3M params |
FineWeb sample (4 files) 0.15× |
5.33 | — | Keep | 1.19873 | 1.21334 | 1.23271 |
| EN0 en0_shared32k_d125m_ctx1024 |
21 Sep 2026 | d125m 110.3M params |
FineWeb sample (4 files) 0.41× |
13.89 | — | Keep | 1.15174 | 1.16785 | 1.18255 |
No runs match these filters. Clear the search or pick “All”.
Cost “—” means no cloud bill (our own GPU at home) or a cost that was not recorded; “Free” is free cloud compute. GPU-h is GPU hours.