en0_shared32k_d125m_ctx1024
d125m (110.3M parameters) on FineWeb sample (4 files), trained on RTX 3060 12GB (Home GPU). Started 21 Sep 2026, ended 21 Sep 2026.
Hypothesis
Our training pipeline, pointed at English web text (1.33B FineWeb tokens, 125M-class model), lands where a standard GPT-2 124M does at the same budget.
What we learned
Yes: FineWeb validation 1.1517. Measured the way the reference was, 1.1781 vs GPT-2 124M at 1.1755 (+0.2%): on the curve. The pipeline reproduces the field.
Scores
Lower is better on the headline.
| Test set | EN0 |
|---|---|
| FineWeb validation headline About 15,000 FineWeb web documents never used in training, scored with a 1,024-token window. |
1.15174 |
| FineWeb validation, clean The same documents minus any that overlap the training text. |
1.16785 |
| Pooled, five sets All five English test sets pooled. |
1.18255 |
| WikiText-103 test Good and featured Wikipedia articles, a standard language-modelling benchmark. |
1.21418 |
| enwik8 test Raw Wikipedia markup; a character-level compression benchmark. |
1.42598 |
| text8 test Lower-cased Wikipedia text with markup stripped; a character-level benchmark. |
1.30127 |
| FineWeb validation, during training The validation number measured during training, the way the GPT-2 124M reference was measured, so the two can be compared. |
1.1781 |
Bits per byte: how many bits the model needs, on average, to predict each byte of text it never saw in training. Lower is better, and it compares models with different tokenizers fairly, because every model is charged for the same bytes.
Curves
Training loss
Cross-entropy per token, by step. The first few percent of the run, far higher, run off the top; hover or the table has every value.
Held-out bits per byte
The validation split, evaluated during training, by step. Lower is better.
Throughput
Tokens per second, by step.
Model
- Preset
- d125m
- Parameters
- 110,316,288
- Outside embeddings
- 84,953,856
- Layers · heads · width
- 12 · 12 · 768
- Tokenizer
- Shared unigram 32k
Data
- Slice
- FineWeb sample (4 files)
- Words
- —
- Training tokens
- 3,253,318,552
- Passes
- 0.41 passes
- Tokens seen
- 1,332,314,112
Compute
- GPU
- RTX 3060 12GB
- Where
- Home GPU
- Steps
- 27,106 / 27,106
- GPU hours
- 13.89
- Cost
- —
- Spot restarts
- —
Lineage
homed125m