MATH0-MID

Math Running

math0_mid

on OpenWebMath + DeepMind maths (70/30), trained on RTX 3060 12GB (Home GPU). .

Hypothesis

250M more pre-training tokens for MATH0, 70% its maths web pages and 30% DeepMind arithmetic, algebra and number drills, teach it the arithmetic fine-tuning could not.

What we learned

Running: continued pre-training from MATH0; scores follow.

Scores

No scores published yet. The held-out scores are published when the run finishes; the live card above shows the latest evaluation during training.

Curves

Training loss

Cross-entropy per token, by step. The first few percent of the run, far higher, run off the top; hover or the table has every value.

Held-out bits per byte

The validation split, evaluated during training, by step. Lower is better.

Throughput

Tokens per second, by step.

Model

Preset
—
Parameters
—
Outside embeddings
—
Layers · heads · width
— · — · —
Tokenizer
Shared unigram 32k

Muon + AdamW at a third of the peak rate, flat then linear decay, from MATH0

Data

Slice
OpenWebMath + DeepMind maths (70/30)
Words
—
Training tokens
—
Passes
—
Tokens seen
172,032,000

Compute

GPU
RTX 3060 12GB
Where
Home GPU
Steps
3,500 / 5,090
GPU hours
1.95
Cost
—
Spot restarts
—

Lineage

home