visiontest_d_s1338_d20m
d20m (27.1M parameters) on plus_clean subset, trained on RTX 3060 12GB (Home GPU). Started 3 Oct 2026, ended 3 Oct 2026.
Hypothesis
Content control: clean e-text of the same works instead of OCR.
What we learned
0.8% worse than A (nine books are a narrower mix), and both OCR versions lose to it: re-OCR text is an OCR tier, not clean text.
Scores
Bits per byte, lower is better; change against the parent run, VISION-OCR A.
ex-Gītā (headline)
0.7979
bits per byte
+0.8%ex-Gītā, clean_v1
—
bits per byte
Pooled, all five sets
0.7918
bits per byte
+0.8%Validation split
0.7791
bits per byte
−0.5%| Test set | VISION-OCR D s2 | VISION-OCR A (parent) | Change |
|---|---|---|---|
| DCS gold (classical) | 0.7729 | 0.7758 | −0.4% |
| Bhagavad-gītā (memorisation) | 0.6328 | 0.6338 | −0.2% |
| Out of domain | 0.8051 | 0.7976 | +0.9% |
| Prose | 0.7621 | 0.756 | +0.8% |
| Vedic (Ṛgveda) | 1.115 | 1.0937 | +1.9% |
Curves
No training curves were published for this run.Its scores above are what it published.
Model
- Preset
- d20m
- Parameters
- 27,073,024
- Outside embeddings
- 18,881,024
- Layers · heads · width
- 6 · 8 · 512
- Tokenizer
- SLP1 unigram 8k
Data
- Slice
- plus_clean subset
- Words
- —
- Training tokens
- 88,880,889
- Passes
- 1 passes
- Tokens seen
- 88,915,968
Compute
- GPU
- RTX 3060 12GB
- Where
- Home GPU
- Steps
- 1,809 / 1,809
- GPU hours
- 0.23
- Cost
- —
- Spot restarts
- —
Lineage
homed20m