A Sanskrit-only language model

Sansar संसार

Type the start of a verse or a sentence. A model trained from scratch on Sanskrit, and nothing else, writes on.

Try it

0 / 300
Settings

About 7 Devanagari characters per token. Temperature 0 always picks the likeliest word; higher is more adventurous. Each press of Continue samples afresh.

The continuation appears here.

Sansar is a base model. It continues text; it does not answer questions, and it makes things up. Do not treat the output as a quotation or a fact. The first continuation after a quiet spell can take up to half a minute while the model wakes up.

Try an opening

    How Sansar reads Sanskrit

    Every model reads text in tokens. Sansar's tokenizer was trained on Sanskrit only, so it needs far fewer pieces per word than tokenizers built for many languages. Type below to see the pieces.

    Each chip is one token as Sansar sees it. The model reads SLP1, a lossless Latin spelling of Devanagari, so a token can end inside a syllable: those pieces are shown in italics, in IAST.

    Tokens for this text

     

    Fewer tokens means more Sanskrit fits in the same context, and each token carries more meaning. Gemma 3 and Llama 3.2 are counted for the example sentences only: their tokenizer files are not bundled with this site.

    Published averages

    Tokens per word over a fixed Sanskrit test set, exactly as published with the runs data. Lower is better.

    TokenizerVocabularyTokens per wordBytes per token
    IndicBERT v2250,0002.319.1
    Sansar SLP1 unigram 8k8,0003.016.98
    Sarvam-168,0963.436.12
    OpenAI o200k_base200,0193.625.8
    DeepSeek-V3128,8154.724.45
    Qwen3151,6697.552.78
    OpenAI cl100k_base100,2778.22.56
    Gemma 3——6.81
    Llama 3.2——4.68
    Qwen3.5——4.66