f9_slp1_uni8k_d350m_plus_all_3x
d350m (318.4M parameters) on plus_all slice (520.4M words, 3 passes), trained on A100 40GB (Google Cloud, spot). Started 3 Oct 2026, ended 4 Oct 2026.
Hypothesis
F7's 350M model on the full plus_all slice (71.5M more leak-checked words) for three passes beats F7, also on the contamination-clean sets.
What we learned
Winner: ex-Gita 0.5547 (3.0% better than F7), clean_v1 0.5773 (2.8% better). Data and passes changed together, so it does not say which helped. Released as sansar-350m v0.2.0.
Scores
Bits per byte, lower is better; change against the parent run, F7.
| Test set | F9 | F7 (parent) | Change |
|---|---|---|---|
| DCS gold (classical) | 0.5687 | 0.5757 | −1.2% |
| Bhagavad-gītā (memorisation) | 0.1367 | 0.1552 | −11.9% |
| Out of domain | 0.5542 | 0.573 | −3.3% |
| Prose | 0.5406 | 0.5577 | −3.1% |
| Vedic (Ṛgveda) | 0.6892 | 0.7052 | −2.3% |
Curves
Training loss
Cross-entropy per token, by step. The first few percent of the run, far higher, run off the top; hover or the table has every value.
Held-out bits per byte
The validation split, evaluated during training, by step. Lower is better.
Throughput
Tokens per second, by step.
Model
- Preset
- d350m
- Parameters
- 318,424,064
- Outside embeddings
- 302,040,064
- Layers · heads · width
- 24 · 16 · 1,024
- Tokenizer
- SLP1 unigram 8k
Data
- Slice
- plus_all slice
- Words
- 520,400,000
- Training tokens
- 1,696,174,507
- Passes
- 3 passes
- Tokens seen
- 5,088,559,104
Compute
- GPU
- A100 40GB
- Where
- Google Cloud (spot)
- Steps
- 103,527 / 103,527
- GPU hours
- 20.53
- Cost
- $26.69
- Spot restarts
- 0
Lineage
gcpa100spotd350m