recipe_wd_t05_muononly_d125m
on plus_all v2 subset (4 passes) (28.3M words), trained on L4 24GB (AWS, spot). .
Hypothesis
The weight-decay gain comes from decaying the Muon matrices alone.
What we learned
—
Scores
No scores published yet.
Curves
Training loss
Cross-entropy per token, by step. The first few percent of the run, far higher, run off the top; hover or the table has every value.
Held-out bits per byte
The validation split, evaluated during training, by step. Lower is better.
Throughput
Tokens per second, by step.
Model
- Preset
- —
- Parameters
- —
- Outside embeddings
- —
- Layers · heads · width
- — · — · —
- Tokenizer
- SLP1 unigram 8k
Data
- Slice
- plus_all v2 subset (4 passes)
- Words
- 28,300,000
- Training tokens
- 91,236,966
- Passes
- —
- Tokens seen
- —
Compute
- GPU
- L4 24GB
- Where
- AWS (spot)
- Steps
- 0 / 7,425
- GPU hours
- —
- Cost
- —
- Spot restarts
- —
Lineage
awsl4spot