About

About Sansar संसार

A family of language models for Sanskrit only, made by MuseMesh, with everything open: weights, corpus, tokenizer and the training log.

Sansar is a family of language models for Sanskrit only, made by MuseMesh. Each model is trained from scratch; none is a fine-tune of an English model. They learn from an open, licence-labelled 252M-word Sanskrit corpus, through a Sanskrit tokenizer of their own: an 8k-token unigram SentencePiece model over SLP1 transliteration, wrapped so you type and read Devanagari.

These are base models. They continue the text you give them. They do not follow instructions, answer questions or refuse anything, and they invent verses, authors and works. Exact recall of famous verses is close to zero, with one exception: the Bhagavad-gītā occurs again and again in the training text, so the models partly memorise it. That is why the Gītā example on the demo sounds so sure of itself. The others show the model composing.

You can type in Devanagari, or in IAST (for example athāto brahmajijñāsā); IAST is converted to Devanagari for the model, and the line under each result is IAST generated from the Devanagari. The training text also holds some Hindi, Marathi and Pali, so the model sometimes drifts.

How the runs are measured

Bits per byte is how many bits of information the model needs, on average, to predict each byte of Sanskrit text it has never seen; lower is better. The bytes are those of the text in Devanagari (UTF-8), so models with different tokenizers, ours and others', are measured against the same yardstick.

The test texts were frozen before any model was trained and kept out of training: 3,000 sentences of the Digital Corpus of Sanskrit gold standard, 2,470 prose passages, 2,000 web and out-of-domain texts, 1,000 accented Ṛgveda pādas, and all 700 verses of the Bhagavad-gītā.

  • ex-Gītā, the headline number: bits per byte pooled over the four sets other than the Gītā. The Gītā is quoted inside commentaries and epics throughout the training text, so its own column measures memorisation rather than skill.
  • clean_v1: the same sets minus every item that could overlap the training text in a way the held-out masking could not see (some sources type the visarga as an ASCII colon). It is stricter and reads a little higher; compare clean_v1 numbers only with clean_v1 numbers.
  • External models on the runs page were scored zero-shot on the same sets with their own tokenizers and a longer context window than ours, and their training data may include some of the public test texts.

About the runs page

The runs page is published straight from the machine that trains the models: a snapshot of every experiment when anything changes, and a live feed about once a minute while a job is running. Numbers are shown exactly as published. Runs on our own GPU at home have no cloud bill, so their cost shows as “—”; that is not the same as free. A live job that has not reported for 15 minutes is marked stale, and one silent for two hours is hidden.

Models, tokenizer and data

How the demo runs

On a small shared CPU, one token at a time. Your text is not stored; only its length is counted. The demo keeps its own key/value cache so that long continuations do not slow down, checked against the reference forward pass of the released model code. The tokenizer box runs Sansar's released tokenizer; the Qwen3 count comes from Qwen's published tokenizer (Apache-2.0).

Licence

Model weights and tokenizer: CC BY-NC 4.0 (non-commercial use, with attribution to “Sansar, Muse Mesh Private Limited”). Code, including this site: Apache-2.0. For other uses, write to kushal@muse-mesh.com.