F10

Sanskrit Running

f10_slp1_uni8k_d700m_plus_all_v2_3x

d700m (704.1M parameters) on plus_all v2 slice (489.9M words, 3 passes (1.93 done)), trained on A100 40GB (Google Cloud, spot). Started 5 Oct 2026.

Hypothesis

A 700M model (twice F9's size) on the filtered plus_all v2 slice, with the weight decay found on 125M proxies, beats F9.

What we learned

Three passes (96,822 steps) on a spot A100. Scored on the held-out sets when training ends.

Scores

No scores published yet. The held-out scores are published when the run finishes; the live card above shows the latest evaluation during training.

Curves

Training loss

Cross-entropy per token, by step. The first few percent of the run, far higher, run off the top; hover or the table has every value.

Held-out bits per byte

The validation split, evaluated during training, by step. Lower is better.

Throughput

Tokens per second, by step.

Model

Preset
d700m
Parameters
704,128,512
Outside embeddings
679,552,512
Layers · heads · width
24 · 12 · 1,536
Tokenizer
SLP1 unigram 8k

Muon + AdamW, rope, qk_norm, relu^2, untied head, weight decay

Data

Slice
plus_all v2 slice
Words
489,850,000
Training tokens
1,586,326,034
Passes
3 passes (1.93 done)
Tokens seen
3,064,627,200

Compute

GPU
A100 40GB
Where
Google Cloud (spot)
Steps
62,350 / 96,822
GPU hours
26.01
Cost
$33.82
Spot restarts
0

Lineage

gcpa100spotd700m