VISION-OCR D s2

Sanskrit DoneKeep

visiontest_d_s1338_d20m

d20m (27.1M parameters) on plus_clean subset, trained on RTX 3060 12GB (Home GPU). Started 3 Oct 2026, ended 3 Oct 2026.

Hypothesis

Content control: clean e-text of the same works instead of OCR.

What we learned

0.8% worse than A (nine books are a narrower mix), and both OCR versions lose to it: re-OCR text is an OCR tier, not clean text.

Scores

Bits per byte, lower is better; change against the parent run, VISION-OCR A.

ex-Gītā (headline)
0.7979
bits per byte
+0.8%
ex-Gītā, clean_v1
—
bits per byte
Pooled, all five sets
0.7918
bits per byte
+0.8%
Validation split
0.7791
bits per byte
−0.5%
Bits per byte on each test set
Test setVISION-OCR D s2VISION-OCR A (parent)Change
DCS gold (classical)0.77290.7758−0.4%
Bhagavad-gītā (memorisation)0.63280.6338−0.2%
Out of domain0.80510.7976+0.9%
Prose0.76210.756+0.8%
Vedic (Ṛgveda)1.1151.0937+1.9%

Curves

No training curves were published for this run.Its scores above are what it published.

Model

Preset
d20m
Parameters
27,073,024
Outside embeddings
18,881,024
Layers · heads · width
6 · 8 · 512
Tokenizer
SLP1 unigram 8k

Muon + AdamW, rope, qk_norm, relu^2, untied head

Data

Slice
plus_clean subset
Words
—
Training tokens
88,880,889
Passes
1 passes
Tokens seen
88,915,968

Compute

GPU
RTX 3060 12GB
Where
Home GPU
Steps
1,809 / 1,809
GPU hours
0.23
Cost
—
Spot restarts
—

Lineage

homed20m