f4_slp1_uni8k_d125m_muon_arch_s2
d125m (97.2M parameters) on rebuild 5 slice without OCR (446.9M words, 0.93 passes), trained on RTX 3060 12GB (Home GPU). Started 30 Sep 2026, ended 30 Sep 2026.
Hypothesis
Seed noise for the new recipe.
What we learned
ex-Gita 0.6190 vs 0.6189: seed noise on the headline number is 0.02%, far below the earlier 0.6% estimate.
Scores
Bits per byte, lower is better; change against the parent run, F4.
ex-Gītā (headline)
0.619
bits per byte
+0.0%ex-Gītā, clean_v1
—
bits per byte
Pooled, all five sets
0.6083
bits per byte
+0.0%Validation split
0.5904
bits per byte
+0.1%| Test set | F4 s2 | F4 (parent) | Change |
|---|---|---|---|
| DCS gold (classical) | 0.6184 | 0.6204 | −0.3% |
| Bhagavad-gītā (memorisation) | 0.3282 | 0.3271 | +0.3% |
| Out of domain | 0.6165 | 0.6158 | +0.1% |
| Prose | 0.5975 | 0.5974 | +0.0% |
| Vedic (Ṛgveda) | 0.9967 | 1.0098 | −1.3% |
Curves
Training loss
Cross-entropy per token, by step. The first few percent of the run, far higher, run off the top; hover or the table has every value.
Held-out bits per byte
The validation split, evaluated during training, by step. Lower is better.
Throughput
Tokens per second, by step.
Model
- Preset
- d125m
- Parameters
- 97,241,856
- Outside embeddings
- 84,953,856
- Layers · heads · width
- 12 · 12 · 768
- Tokenizer
- SLP1 unigram 8k
Data
- Slice
- rebuild 5 slice without OCR
- Words
- 446,900,000
- Training tokens
- 1,438,157,255
- Passes
- 0.93 passes
- Tokens seen
- 1,332,314,112
Compute
- GPU
- RTX 3060 12GB
- Where
- Home GPU
- Steps
- 27,106 / 27,106
- GPU hours
- 12.38
- Cost
- —
- Spot restarts
- —
Lineage
homed125m