e8_deva_uni8k_syms
d20m (23.2M parameters) on E8 screening slice, trained on RTX 3060 12GB (Home GPU). Started 11 Sep 2026, ended 11 Sep 2026.
Hypothesis
Unigram beats BPE at equal vocabulary, as the tokenizer literature suggests.
What we learned
Yes: pooled 0.7415, the best subword arm and about 2.8% ahead of BPE. Unigram became the tokenizer algorithm.
Scores
Bits per byte, lower is better.
ex-Gītā (headline)
0.7505
bits per byte
ex-Gītā, clean_v1
—
bits per byte
Pooled, all five sets
0.7415
bits per byte
Validation split
0.6781
bits per byte
| Test set | E8-deva-uni |
|---|---|
| DCS gold (classical) | 0.7113 |
| Bhagavad-gītā (memorisation) | 0.5046 |
| Out of domain | 0.7528 |
| Prose | 0.718 |
| Vedic (Ṛgveda) | 1.2775 |
Curves
Training loss
Cross-entropy per token, by step. The first few percent of the run, far higher, run off the top; hover or the table has every value.
Held-out bits per byte
The validation split, evaluated during training, by step. Lower is better.
Throughput
Tokens per second, by step.
Model
- Preset
- d20m
- Parameters
- 23,239,168
- Outside embeddings
- 18,881,024
- Layers · heads · width
- 6 · 8 · 512
- Tokenizer
- Devanagari unigram 8k
Data
- Slice
- E8 screening slice
- Words
- —
- Training tokens
- 639,191,403
- Passes
- 0.59 passes
- Tokens seen
- 375,275,520
Compute
- GPU
- RTX 3060 12GB
- Where
- Home GPU
- Steps
- 7,635 / 7,635
- GPU hours
- 0.84
- Cost
- —
- Spot restarts
- —
homed20m