f4_slp1_uni8k_d125m_muon_arch
d125m (97.2M parameters) on rebuild 5 slice without OCR (446.9M words, 0.93 passes), trained on RTX 3060 12GB (Home GPU). Started 23 Sep 2026, ended 23 Sep 2026.
Hypothesis
The optimiser and architecture recipe that won on English (Muon, rotary positions, QK-norm, ReLU², untied head) transfers to Sanskrit.
What we learned
Yes: ex-Gita 0.6189, 1.9% better than F2-noocr on identical data and steps, Vedic 6.8% better. About 40% of its English gain.
Scores
Bits per byte, lower is better; change against the parent run, F2-noocr.
| Test set | F4 | F2-noocr (parent) | Change |
|---|---|---|---|
| DCS gold (classical) | 0.6204 | 0.6319 | −1.8% |
| Bhagavad-gītā (memorisation) | 0.3271 | 0.3545 | −7.7% |
| Out of domain | 0.6158 | 0.626 | −1.6% |
| Prose | 0.5974 | 0.609 | −1.9% |
| Vedic (Ṛgveda) | 1.0098 | 1.0836 | −6.8% |
Curves
Training loss
Cross-entropy per token, by step. The first few percent of the run, far higher, run off the top; hover or the table has every value.
Held-out bits per byte
The validation split, evaluated during training, by step. Lower is better.
Throughput
Tokens per second, by step.
Model
- Preset
- d125m
- Parameters
- 97,241,856
- Outside embeddings
- 84,953,856
- Layers · heads · width
- 12 · 12 · 768
- Tokenizer
- SLP1 unigram 8k
Data
- Slice
- rebuild 5 slice without OCR
- Words
- 446,900,000
- Training tokens
- 1,438,157,255
- Passes
- 0.93 passes
- Tokens seen
- 1,332,314,112
Compute
- GPU
- RTX 3060 12GB
- Where
- Home GPU
- Steps
- 27,106 / 27,106
- GPU hours
- 12.32
- Cost
- —
- Spot restarts
- —
Lineage
homed125m