Sansar 700M is a Sanskrit-only language model with 704 million parameters, trained from scratch on one rented A100 in about 40 hours, for about $52 of spot time. On our held-out Sanskrit sets it scores better than Gemma 3 4B, Qwen3-4B, Llama 3.2 3B and Sarvam-1 — with 0.7B parameters against their 2.5B to 4B. The weights, the corpus and the tokenizer are public, and so is every run that led here.
Held-out Sanskrit, bits per byte (lower is better)
- Sansar 700M0.70B params0.531
- Sansar 350M0.32B params0.555
- Gemma 3 4B3.88B params0.697
- Qwen3-4B4.02B params0.707
- Llama 3.2 3B3.21B params0.712
- Sarvam-12.53B params0.747
The model is on Hugging Face at MuseMesh/sansar-700m, you can try it in the browser at mume.ai/sansar, and the training log — every run we did, finished, failed or still going, with what it cost — is at mume.ai/lab. This post is what we built, how it compares, and where we were wrong, including a hole in our own test set.
Why a model that only speaks Sanskrit
We wanted to see what happens when the data, the tokenizer and the evaluation are all built for one language, at a size a small team can actually train. Sanskrit suits that experiment. The body of text is finite, so we could audit it source by source instead of crawling and hoping. And general models handle it badly, starting before the model sees a single token — in the tokenizer.
None of our models starts from an English checkpoint. The three smallest were trained on one RTX 3060 at home; the two largest on a rented A100.
Every released Sansar model, bits per byte (lower is better)
- 0.718
- 0.643*
- 0.604
- 0.555
- 0.531
| Model | Parameters | Trained on | Held-out | clean_v1 |
|---|---|---|---|---|
| sansar-20m | 27M | RTX 3060, 51 min | 0.7177 | — |
| sansar-60m | 63M | RTX 3060, 7.5 h | 0.6434* | — |
| sansar-125m | 97M | RTX 3060, 24.7 h | 0.6039 | 0.6254 |
| sansar-350m | 318M | A100 spot, 20.5 h, ~$27 | 0.5547 | 0.5773 |
| sansar-700m | 704M | A100 spot, 40.3 h, ~$52 | 0.5307 | 0.5541 |
What the number means
Bits per byte is how many bits of surprise a model needs, on average, for each byte of Devanagari text it reads. Lower is better. Because it is counted per byte of text rather than per token, it compares models with completely different tokenizers fairly — which matters here, because the tokenizers differ by a factor of almost three.
The held-out sets are frozen, excluded from training by key, and masked out inside training records: classical sentences from the DCS gold corpus, prose, Vedic (accented Ṛgveda pādas) and out-of-domain web text, pooled. We leave the Bhagavad-gītā out of the headline number, because it is quoted all over the training text and its score measures memorisation rather than Sanskrit.
How it compares
We scored well-known open base models on the same held-out sets through the same code: the same items, the same cleaning, the same Devanagari byte count. Each reads the text with its own tokenizer.
| Model | Params | Held-out | clean_v1 | Classical | Web | Prose | Vedic |
|---|---|---|---|---|---|---|---|
| Sansar 700M | 0.70B | 0.5307 | 0.5541 | 0.5513 | 0.5286 | 0.5186 | 0.6670 |
| Sansar 350M | 0.32B | 0.5547 | 0.5773 | 0.5686 | 0.5543 | 0.5406 | 0.6893 |
| Gemma 3 4B pt | 3.88B | 0.6965 | 0.7365 | 0.9671 | 0.6622 | 0.6572 | 1.2084 |
| Qwen3-4B-Base | 4.02B | 0.7071 | 0.7450 | 0.9639 | 0.6688 | 0.6816 | 1.2543 |
| Qwen3.5-4B-Base | 4.21B | 0.7114 | 0.7569 | 0.9883 | 0.6610 | 0.6873 | 1.6024 |
| Llama 3.2 3B | 3.21B | 0.7122 | 0.7607 | 1.0386 | 0.6653 | 0.6784 | 1.3619 |
| Sarvam-1 | 2.53B | 0.7465 | 0.8166 | 1.1287 | 0.6737 | 0.7473 | 1.6491 |
Sansar 700M needs 24% fewer bits per byte than the best of them, Gemma 3 4B (0.5307 against 0.6965; 25% on clean_v1), with 18% of its parameters. The general models do best on web text and worst on classical and Vedic Sanskrit, which is where the gap is widest.
What is not in the table: Krutrim-2 (12B). In an earlier pass in September — 4-bit weights, verse references stripped — the instruction-tuned Krutrim-2 scored about 0.512, better than our 0.5307. That pass is not strictly comparable and we have not re-run it. So this is not “the best Sanskrit model”. The claim is narrower: on our held-out Sanskrit sets, Sansar 700M beats Gemma 3 4B, Qwen3-4B, Llama 3.2 3B and Sarvam-1.
Three more caveats, so the table is read the way it was measured:
- Tokenizers differ. That is what bits per byte is designed to neutralise, and it does.
- The external models had more context. They were scored zero-shot with a 2,048-token window — four times our 512 — and with the start-of-text token in every window where the model has one.
- Their training data may contain our test texts. Public prose and web pages are exactly what general models are trained on.
The table says what a Sanskrit-only model buys at this size. It does not say which model to pick for your task: ours is a base model that cannot answer a question. The full table, including smaller models, is on the lab.
A tokenizer that does not cut letters in half
Tokens per word of classical Sanskrit (DCS gold sentences), measured by the same harness for each tokenizer:
| Tokenizer | Vocabulary | Tokens per word |
|---|---|---|
| Sansar 8k | 8,000 | 2.82 |
| Sarvam-1 | 68,096 | 3.57 |
| GPT-4o (o200k_base) | 200,019 | 3.91 |
| DeepSeek-V3 | 128,815 | 4.98 |
| Qwen3 | 151,669 | 7.84 |
A vocabulary nineteen times smaller than Qwen3’s, and under half the tokens per word. More to the point, the Sansar tokenizer never separates a vowel sign from its consonant: 0 orphaned vowel signs per 1,000 tokens, against 94 to 464 for the public tokenizers we tested. And it is lossless — all 9,170 held-out texts decode back to exactly the text that went in.
It is a unigram SentencePiece model trained on SLP1, a one-letter-per-sound ASCII encoding of Devanagari from the Sanskrit Library, with a wrapper that takes and returns Devanagari. In matched 20M models, unigram beat BPE by 2.8% on bits per byte. The weak spot is Vedic: accent marks fragment the pieces, to 3.85 tokens per word.
A corpus where every record carries its licence
The Sansar Sanskrit Corpus has 252,303,814 words in 1,829,785 records from 21 sources, all in Devanagari. We only included sources whose licence allows redistribution. Every record carries its source, an SPDX licence id, a URL and a one-line credit, and there is one config per licence, so you can load exactly the licences you are able to use:
from datasets import load_dataset
by = load_dataset(
"MuseMesh/sansar-sanskrit-corpus", "cc-by-4.0",
split="train", revision="v0.1.0",
)Two things to know before you use it:
- 72% of the words are OCR. They come from Sangraha (AI4Bharat), which is OCR of digitised books. In our audit, 2 of 10 random Sangraha chunks were visibly garbled. If quality matters more than size, use the curated configs.
- The models were not trained on this corpus alone. Their training mix also includes non-commercial and no-licence-stated text we cannot redistribute. The corpus is the openly licensed part, and that mix is why the weights are non-commercial.
What changed from 350M to 700M
Sansar 700M is 4.3% better than the 350M on the standard held-out sets (0.5307 against 0.5547), 4.0% better on clean_v1 (0.5541 against 0.5773), and better on every individual set. It saw 4.76 billion tokens: three passes over a filtered 490-million-word training slice.
Three things changed at once: the model size (318M to 704M), the training slice (quality filters, plus 1.9M words of open web crawls), and weight decay. We did not run them separately, so the 4% is not split between them, and we are not going to pretend to know the split.
Our test set leaked through a colon
This is the part we would most like other people building small-language corpora to steal.
Some of our training text types the visarga — ः, the little two-dot mark — as an ASCII colon. Our overlap gate and our training-time mask both ignored colons. So held-out passages written that way looked different to the gate, matched nothing, and slipped into training.
Once we found it we re-scanned everything with a colon-aware rule and built clean_v1: the held-out sets minus every affected item, 900 of 9,170. Most hits were one stock phrase of 16 to 23 characters; real near-copies were about 94 items. The damage turned out to be small but not zero. The 350M’s gain over its previous version, 3.03% on the standard sets, became 2.77% on clean_v1. The 700M’s 4.3% over the 350M becomes 4.0%.
Every model card from 125M up now reports both columns, and the public corpus was cleaned with the colon-aware rule. If you only read one column, read clean_v1.
Two other things the runs taught us
Old OCR text hurts in proportion to its share. We trained 125M models with archive.org OCR at 0%, 12.5% and 45% of the tokens. Held-out bits per byte came out 0.6306, 0.6336 and 0.6399 — almost exactly a straight line, about a third of a percent worse per 10 points of OCR, with no threshold and no sweet spot above zero. These effects are small and measured at small scale, and we have not ablated Sangraha itself.
Weight decay beat the architecture tricks. On 125M proxies, decoupled weight decay on the Muon matrices cut held-out bits per byte by 3.3% and Vedic by 11%, against a seed noise bar of 0.42%. Value embeddings, a logit cap and document masking bought nothing. The 700M was trained with this decay — but, as above, alongside two other changes, so we have not shown its effect at 700M on its own.
What it is not
These are base models. They continue Sanskrit text; they do not follow instructions or answer questions, and they make things up. Exact verse recall on our 600-verse completion test is 4 of 600.
Bigger models also memorise more of what they see often. The Bhagavad-gītā costs the 700M 0.076 bits per byte, against 0.137 for the 350M — it reproduces much of it nearly verbatim. That is why the Gītā stays out of every headline number in this post.
Try it
The quickest way is the browser: mume.ai/sansar now runs Sansar 700M, with nothing to install. Give it the start of a line and it continues it.
Or locally, with Transformers:
from transformers import AutoModelForCausalLM, AutoTokenizer
repo = "MuseMesh/sansar-700m"
tok = AutoTokenizer.from_pretrained(repo, revision="v0.1.0", trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(repo, revision="v0.1.0", trust_remote_code=True)
inputs = tok("संस्कृतं नाम दैवी वाक्", return_tensors="pt")
out = model.generate(**inputs, max_new_tokens=50, do_sample=True, temperature=0.8, top_k=100)
print(tok.decode(out[0], skip_special_tokens=True))trust_remote_code=True runs three small files from the repository, so read them first. The context is 512 tokens and the reference implementation has no key/value cache, so long generations are slow. The weights are 1.4 GB in bfloat16.
What is next
The 700M saw its training slice three times, and the passes are the open question. On its own validation split, bits per byte fell from 0.541 after the first pass to 0.489 after the second and 0.449 after the third — but the gap between training and validation loss widened from 0.15 to 0.31 nats per token over the same passes. That split is not the frozen held-out set, so those numbers do not compare with the ones above; what they say is that each pass still helps and the model is memorising more as it goes. A proxy test of three, four and five passes runs tonight, and the answer decides the shape of the next run. It will show up on the lab as it happens.
Beyond that: an OCR model for Sanskrit pages, and a commercially licensed version trained only on permissively licensed text, as a separate release.
If you work with Sanskrit, tell us what is wrong — bad sources, errors in the output, missing licences, evaluations you would trust more than ours. The Hugging Face discussion tabs are open, and corrections or takedown requests for the corpus go to kushal@muse-mesh.com.
Licence
Weights and tokenizer are CC BY-NC 4.0: non-commercial use, with attribution to “Sansar, Muse Mesh Private Limited”. Code in the model repositories is Apache-2.0. Corpus records keep the licence of their source, and our compilation is CC BY 4.0. For commercial use, write to kushal@muse-mesh.com. Everything is under huggingface.co/MuseMesh.