math0_mid
on OpenWebMath + DeepMind maths (70/30), trained on RTX 3060 12GB (Home GPU). .
Hypothesis
250M more pre-training tokens for MATH0, 70% its maths web pages and 30% DeepMind arithmetic, algebra and number drills, teach it the arithmetic fine-tuning could not.
What we learned
Running: continued pre-training from MATH0; scores follow.
Scores
No scores published yet. The held-out scores are published when the run finishes; the live card above shows the latest evaluation during training.
Curves
Training loss
Cross-entropy per token, by step. The first few percent of the run, far higher, run off the top; hover or the table has every value.
Held-out bits per byte
The validation split, evaluated during training, by step. Lower is better.
Throughput
Tokens per second, by step.
Model
- Preset
- —
- Parameters
- —
- Outside embeddings
- —
- Layers · heads · width
- — · — · —
- Tokenizer
- Shared unigram 32k
Data
- Slice
- OpenWebMath + DeepMind maths (70/30)
- Words
- —
- Training tokens
- —
- Passes
- —
- Tokens seen
- 172,032,000
Compute
- GPU
- RTX 3060 12GB
- Where
- Home GPU
- Steps
- 3,500 / 5,090
- GPU hours
- 1.95
- Cost
- —
- Spot restarts
- —
Lineage
home