Sanskrit training runs

Sansar: language models for Sanskrit only, trained from scratch on an open 252M-word corpus with a Sanskrit tokenizer of their own.

The headline score is ex-Gītā. Bits per byte pooled over four held-out Sanskrit sets: classical prose and verse, web text and the Ṛgveda. The Bhagavad-gītā is left out because the models partly memorise it.

Bits per byte: how many bits the model needs, on average, to predict each byte of text it never saw in training. Lower is better, and it compares models with different tokenizers fairly, because every model is charged for the same bytes.

Try the Sansar demo · how Sanskrit is measured · the Sansar runs page

Every score on this track
ex-Gītā (lower is better)
Bits per byte pooled over four held-out Sanskrit sets: classical prose and verse, web text and the Ṛgveda. The Bhagavad-gītā is left out because the models partly memorise it.
ex-Gītā, clean_v1 (lower is better)
The same four sets minus every item that could overlap the training text. Stricter; compare only with other clean_v1 numbers.
Pooled, five sets (lower is better)
All five held-out sets pooled, the Gītā included.
Validation split (lower is better)
A slice of the training corpus held back from training, scored during the run.

Live now

Progress

Lower is better on the headline.

Best ex-Gītā over time

Each dot is a scored run, placed at the day it ended; the line is the best score so far. Numbered lines are this track's milestones, listed below.

Milestones
  1. 11 Sep 2026 · First end-to-end training run on the home GPU
  2. 12 Sep 2026 · Tokenizer frozen: SLP1 unigram 8k
  3. 13 Sep 2026 · F0: first model on the full corpus (60M)
  4. 14 Sep 2026 · Held-out masking: the honest baseline F1b
  5. 20 Sep 2026 · OCR text hurts in proportion to its share
  6. 24 Sep 2026 · F4: the Muon + architecture recipe transfers to Sanskrit
  7. 1 Oct 2026 · First release: the Sansar tokenizer on Hugging Face
  8. 3 Oct 2026 · F7: first 350M model, first cloud run
  9. 4 Oct 2026 · F9: best model so far (ex-Gita 0.5547)
  10. 4 Oct 2026 · Models (20M to 350M) and a 252M-word corpus public on Hugging Face
  11. 5 Oct 2026 · F10 (700M) starts training
  12. 6 Oct 2026 · Sansar 350m compared with open base models up to 4B

Model size and ex-Gītā

Parameters (log scale) against ex-Gītā in bits per byte: this track's runs and other models measured the same way. Hover a point for its name.

Reference models

Other models measured on the same sets, for comparison. Notes say how they were measured.

Reference models scored on the Sanskrit track's test sets
ModelParametersex-Gītāex-Gītā, clean_v1
Krutrim-2 12B (instruct) Scored in an earlier pass (4-bit weights, verse references stripped), so not strictly comparable; it beats Sansar 350m there.—0.512—
Gemma 3 4B pt3.9B0.69650.7365
Qwen3-4B-Base4B0.70710.745
Qwen3.5-4B-Base4.2B0.71140.7569
Llama 3.2 3B3.2B0.71220.7607
Sarvam-12.5B0.74650.8166
Qwen3-1.7B-Base1.7B0.78720.8235
Llama 3.2 1B1.2B0.80720.8581
Qwen3.5-2B-Base1.9B0.81170.8609
Gemma 3 1B pt999.9M0.83990.8778
Qwen3-0.6B-Base596M0.89920.9321

All Sanskrit runs

Published · 69 runs

All Sanskrit runs. Column headers sort the table.
Data
RECIPE WD τ0.5+MIX
recipe_wd_t05_mix_d125m
6 Oct 2026 d125m
97.2M params
plus_all v2 subset (4 passes)
28.3M words · 3.51×
2.2 $2.03 — — — —
F10 Running
f10_slp1_uni8k_d700m_plus_all_v2_3x
5 Oct 2026 d700m
704.1M params
plus_all v2 slice
489.9M words · 1.92×
25.9 $33.67 — — — —
RECIPE MIX
recipe_mix_d125m
5 Oct 2026 d125m
97.2M params
plus_all v2 subset (4 passes)
28.3M words · 3.51×
3.38 — Discard 0.7236 0.7532 0.714
RECIPE ARCH
recipe_arch_d125m
5 Oct 2026 d125m
103.4M params
plus_all v2 subset (4 passes)
28.3M words · 4.02×
3.46 — Discard 0.7542 0.804 0.7473
RECIPE R0docs
recipe_r0docs_d125m
5 Oct 2026 d125m
97.2M params
plus_all v2 subset (4 passes)
28.3M words · 4.02×
3.38 — Neutral 0.7262 0.7607 0.7167
RECIPE WD τ0.25
recipe_wd_t025_d125m
4 Oct 2026 d125m
97.2M params
plus_all v2 subset (4 passes)
28.3M words · 4×
3.39 — Keep 0.6946 0.7178 0.6865
RECIPE WD τ0.5
recipe_wd_t05_d125m
4 Oct 2026 d125m
97.2M params
plus_all v2 subset (4 passes)
28.3M words · 4×
3.39 — Winner 0.691 0.7149 0.6822
RECIPE WD τ1
recipe_wd_t1_d125m
4 Oct 2026 d125m
97.2M params
plus_all v2 subset (4 passes)
28.3M words · 4×
3.39 — Keep 0.6971 0.7225 0.6877
RECIPE R0 s2
recipe_r0_s2_d125m
4 Oct 2026 d125m
97.2M params
plus_all v2 subset (4 passes)
28.3M words · 4×
3.48 — Keep 0.7136 0.7402 0.7037
RECIPE R0
recipe_r0_d125m
4 Oct 2026 d125m
97.2M params
plus_all v2 subset (4 passes)
28.3M words · 4×
3.39 — Keep 0.7152 0.7427 0.7052
F10-AB b
f10ab_b_d60m
3 Oct 2026 d60m
68.9M params
plus_all band, filtered
90M words · 1.07×
1.91 — Neutral 0.6814 — 0.6735
F10-AB a
f10ab_a_d60m
3 Oct 2026 d60m
68.9M params
plus_all band
96M words · 1×
1.91 — Neutral 0.6831 — 0.6754
E-CTX base s2
i9small_base_s2_d60m
3 Oct 2026 d60m
68.9M params
plus_clean subset
1×
1.99 — Keep 0.6842 — 0.6761
E-VOCAB 16k
i9small_vocab16k_d60m
3 Oct 2026 d60m
81.2M params
plus_clean subset
1×
1.95 — Neutral 0.6843 — 0.6757
E-VEDIC 3x
i9small_vedic3x_d60m
3 Oct 2026 d60m
68.9M params
plus_clean subset + Vedic x3
0.99×
1.99 — Keep 0.683 — 0.675
E-CTX 1024
i9small_ctx1024_d60m
3 Oct 2026 d60m
68.9M params
plus_clean subset
1×
2.08 — Neutral 0.6856 — 0.6778
F9
f9_slp1_uni8k_d350m_plus_all_3x
3 Oct 2026 d350m
318.4M params
plus_all slice
520.4M words · 3×
20.53 $26.69 Winner 0.5547 0.5773 0.5394
E-CTX base
i9small_base_d60m
3 Oct 2026 d60m
68.9M params
plus_clean subset
1×
2.03 — Keep 0.6825 — 0.6744
VISION-OCR D s2
visiontest_d_s1338_d20m
3 Oct 2026 d20m
27.1M params
plus_clean subset
1×
0.23 — Keep 0.7979 — 0.7918
VISION-OCR D
visiontest_d_s1337_d20m
3 Oct 2026 d20m
27.1M params
plus_clean subset
1×
0.23 — Keep 0.7993 — 0.7932
VISION-OCR C s2
visiontest_c_s1338_d20m
3 Oct 2026 d20m
27.1M params
plus_clean subset
1×
0.23 — Neutral 0.8002 — 0.7945
VISION-OCR B s2
visiontest_b_s1338_d20m
3 Oct 2026 d20m
27.1M params
plus_clean subset
1×
0.23 — Neutral 0.8004 — 0.7946
VISION-OCR A s2
visiontest_a_s1338_d20m
3 Oct 2026 d20m
27.1M params
plus_clean subset
1×
0.23 — Keep 0.7937 — 0.7878
VISION-OCR C
visiontest_c_s1337_d20m
3 Oct 2026 d20m
27.1M params
plus_clean subset
1×
0.23 — Neutral 0.8015 — 0.7958
VISION-OCR B
visiontest_b_s1337_d20m
3 Oct 2026 d20m
27.1M params
plus_clean subset
1×
0.22 — Neutral 0.8033 — 0.7975
VISION-OCR A
visiontest_a_s1337_d20m
3 Oct 2026 d20m
27.1M params
plus_clean subset
1×
0.23 — Keep 0.7913 — 0.7856
F8-small
f5_slp1_uni8k_d125m_plus_t10
2 Oct 2026 d125m
97.2M params
plus_clean + additions
509.3M words · 0.8×
12.3 — Keep 0.6135 — 0.6036
TOK-v3 v3data
tokv3_v3data_d20m
2 Oct 2026 d20m
27.1M params
plus_clean subset
1×
0.82 — Discard 0.7191 — 0.7118
TOK-v3 v3acc
tokv3_v3acc_d20m
2 Oct 2026 d20m
27.1M params
plus_clean subset
1×
0.86 — Discard 0.7187 — 0.7113
TOK-v3 v0.1
tokv3_v01_d20m
2 Oct 2026 d20m
27.1M params
plus_clean subset
1×
0.85 — Keep 0.7177 — 0.7102
F7
f7_slp1_uni8k_d350m_plus_clean_2x
2 Oct 2026 d350m
318.4M params
plus_clean slice
448.9M words · 1.84×
14.01 $18.22 Keep 0.572 0.5937 0.5568
F6-clean-2x
f5_slp1_uni8k_d125m_plus_clean_2x
1 Oct 2026 d125m
97.2M params
plus_clean slice
448.9M words · 1.84×
24.74 — Keep 0.6039 0.6254 0.592
F6-clean
f5_slp1_uni8k_d125m_plus_clean
30 Sep 2026 d125m
97.2M params
plus_clean slice
448.9M words · 0.92×
12.35 — Keep 0.6145 — 0.6039
F4 s2
f4_slp1_uni8k_d125m_muon_arch_s2
30 Sep 2026 d125m
97.2M params
rebuild 5 slice without OCR
446.9M words · 0.93×
12.38 — — 0.619 — 0.6083
F6
f5_slp1_uni8k_d125m_plus
29 Sep 2026 d125m
97.2M params
plus slice 12.32 — Discard 0.6142 — 0.6037
F5-filtered v2
f5_slp1_uni8k_d125m_filtered_v2
27 Sep 2026 d125m
97.2M params
line-filtered slice v2 12.3 — Discard 0.6256 — 0.6146
F5-curated
f5_slp1_uni8k_d125m_curated
26 Sep 2026 d125m
97.2M params
curated e-text subset
173M words
12.34 — Discard 0.6575 — 0.6465
F5-filtered
f5_slp1_uni8k_d125m_filtered
26 Sep 2026 d125m
97.2M params
line-filtered slice
355M words
12.28 — Discard 0.6238 — 0.6128
F4
f4_slp1_uni8k_d125m_muon_arch
23 Sep 2026 d125m
97.2M params
rebuild 5 slice without OCR
446.9M words · 0.93×
12.32 — Keep 0.6189 — 0.6082
F3-mix12
f3_slp1_uni8k_d125m_mix12
19 Sep 2026 d125m
91.5M params
rebuild 5 slice, 12.5% OCR
505.8M words · 0.81×
10.89 — Keep 0.6336 — 0.6235
F2-noocr
f2_slp1_uni8k_d125m_noocr
19 Sep 2026 d125m
91.5M params
rebuild 5 slice without OCR
446.9M words · 0.93×
10.97 — Keep 0.6306 — 0.6205
F2
f2_slp1_uni8k_d125m
18 Sep 2026 d125m
91.5M params
rebuild 5 slice
798.3M words · 0.5×
10.95 — Keep 0.6399 — 0.6296
F1b s2
f1b_slp1_uni8k_d125m_s2
15 Sep 2026 d125m
91.5M params
rebuild 3 slice
414M words · 1×
10.84 — Keep 0.641 — 0.6308
F1b
f1b_slp1_uni8k_d125m
14 Sep 2026 d125m
91.5M params
rebuild 3 slice
414M words · 1×
10.75 — Keep 0.6372 — 0.6271
F1
f1_slp1_uni8k_d125m
13 Sep 2026 d125m
91.5M params
rebuild 3 slice
414M words · 1×
10.78 — Keep 0.6273 — 0.6132
F0
f0_slp1_uni8k_d60m
13 Sep 2026 d60m
63.2M params
rebuild 3 slice
414M words · 1×
7.49 — Keep 0.6434 — 0.6282
E15-char-1024
e15_slp1_char_ctx1024
12 Sep 2026 d20m
19.6M params
E8 screening slice
0.59×
2.12 — Keep 0.7609 — 0.7528
E9-16k s2
kg_e9s2_slp1_uni16k
12 Sep 2026 d20m
24M params
E8 screening slice
0.59×
1.62 Free Keep 0.766 — 0.7576
E9-8k s2
kg_e9s2_slp1_uni8k
12 Sep 2026 d20m
22.7M params
E8 screening slice
0.59×
1.69 Free Keep 0.7681 — 0.7597
Gemma 4 E2B CPT Failed
kg_gemma4_e2b_cpt30m
12 Sep 2026 —
4B params
Devanagari slice 8.07 Free Discard 2.6118 — 2.6143
E14-8k
e14_slp1_uni8k
12 Sep 2026 d60m
63.2M params
E8 screening slice
0.59×
2.13 — Keep 0.7117 — 0.7016
E14-16k
e14_slp1_uni16k
12 Sep 2026 d60m
69.3M params
E8 screening slice
0.59×
2.1 — Keep 0.7134 — 0.703
E11-bpe-dropout
kg_e11_slp1_bpe8k_dropout
12 Sep 2026 d20m
23.2M params
E8 screening slice
0.59×
1.65 Free Discard 0.814 — 0.8059
E11-uni-sampling
kg_e11_slp1_uni8k_sampling
12 Sep 2026 d20m
23.2M params
E8 screening slice
0.59×
2.03 Free Discard 0.9526 — 0.9455
E9-32k
e9_slp1_uni32k
12 Sep 2026 d20m
23.1M params
E8 screening slice
0.59×
0.76 — Discard 0.7671 — 0.7576
E9-16k
e9_slp1_uni16k
12 Sep 2026 d20m
24M params
E8 screening slice
0.59×
0.91 — Keep 0.7556 — 0.7465
E10-bpe16k
kg_slp1_bpe16k
12 Sep 2026 d20m
27.3M params
E8 screening slice
0.59×
1.28 Free Discard 0.7762 — 0.7672
E9-8k
e9_slp1_uni8k
12 Sep 2026 d20m
22.7M params
E8 screening slice
0.59×
0.95 — Keep 0.7552 — 0.7466
E10-uni16k
kg_slp1_uni16k_d20m
12 Sep 2026 d20m
27.3M params
E8 screening slice
0.59×
1.41 Free Keep 0.7626 — 0.7536
E9-4k
e9_slp1_uni4k
12 Sep 2026 d20m
23M params
E8 screening slice
0.59×
1.07 — Discard 0.7581 — 0.7497
E8-slp1-uni
e8_slp1_uni8k
11 Sep 2026 d20m
23.2M params
E8 screening slice
0.59×
0.88 — Keep 0.7537 — 0.745
E8-slp1-char
e8_slp1_char
11 Sep 2026 d20m
19.4M params
E8 screening slice
0.59×
1.85 — Keep 0.7503 — 0.7419
E8-slp1-bpe
e8_slp1_bpe8k
11 Sep 2026 d20m
23.2M params
E8 screening slice
0.59×
0.78 — Discard 0.7707 — 0.7619
E8-deva-uni
e8_deva_uni8k_syms
11 Sep 2026 d20m
23.2M params
E8 screening slice
0.59×
0.84 — Winner 0.7505 — 0.7415
E8-deva-bpe
e8_deva_bpe8k_nosyms
11 Sep 2026 d20m
23.2M params
E8 screening slice
0.59×
0.77 — Discard 0.7713 — 0.7624
E8-deva-bpe-sym
e8_deva_bpe8k_syms
11 Sep 2026 d20m
23.2M params
E8 screening slice
0.59×
0.78 — Discard 0.7719 — 0.7629
Smoke test
smoke_d20m
11 Sep 2026 d20m
23.2M params
early test slice 1 — — — — —
RECIPE WD τ0.5 s2 Queued
recipe_wd_t05_s2_d125m
— — plus_all v2 subset (4 passes)
28.3M words
— — — — — —
RECIPE WD τ0.5 muon-only Queued
recipe_wd_t05_muononly_d125m
— — plus_all v2 subset (4 passes)
28.3M words
— — — — — —

Cost “—” means no cloud bill (our own GPU at home) or a cost that was not recorded; “Free” is free cloud compute. GPU-h is GPU hours.