A Sanskrit-only language model
Sansar संसार
Type the start of a verse or a sentence. A model trained from scratch on Sanskrit, and nothing else, writes on.
Try it
The continuation appears here.
Sansar is a base model. It continues text; it does not answer questions, and it makes things up. Do not treat the output as a quotation or a fact. The first continuation after a quiet spell can take up to half a minute while the model wakes up.
Try an opening
How Sansar reads Sanskrit
Every model reads text in tokens. Sansar's tokenizer was trained on Sanskrit only, so it needs far fewer pieces per word than tokenizers built for many languages. Type below to see the pieces.
Tokens for this text
Published averages
| Tokenizer | Vocabulary | Tokens per word | Bytes per token |
|---|---|---|---|
| IndicBERT v2 | 250,000 | 2.31 | 9.1 |
| Sansar SLP1 unigram 8k | 8,000 | 3.01 | 6.98 |
| Sarvam-1 | 68,096 | 3.43 | 6.12 |
| OpenAI o200k_base | 200,019 | 3.62 | 5.8 |
| DeepSeek-V3 | 128,815 | 4.72 | 4.45 |
| Qwen3 | 151,669 | 7.55 | 2.78 |
| OpenAI cl100k_base | 100,277 | 8.2 | 2.56 |
| Gemma 3 | — | — | 6.81 |
| Llama 3.2 | — | — | 4.68 |
| Qwen3.5 | — | — | 4.66 |