OCR training runs

Reading scanned Sanskrit books: an open vision-language model fine-tuned for page OCR, and the bulk OCR jobs that turn scans into corpus text.

The headline score is Median letter error rate. The share of letters read wrong on 54 held-out book pages, median over pages: 0.0096 means 0.96% of letters.

Letter error rate: the share of letters an OCR model reads wrong, as a fraction (0.0096 means 0.96% of letters). Lower is better.

Live now

Progress

Lower is better on the headline.

Best Median letter error rate over time

Each dot is a scored run, placed at the day it ended; the line is the best score so far. Numbered lines are this track's milestones, listed below.

Milestones
  1. 6 Oct 2026 · OCR-VLM v1 beats Google Vision on Sanskrit pages

Model size and Median letter error rate

Parameters (log scale) against Median letter error rate: this track's runs. Hover a point for its name.

All OCR runs

Published · 4 runs

All OCR runs. Column headers sort the table.
Data
OCR pilot test
ocr_pilot100k_test
6 Oct 2026 — — 0.57 $0.52 — —
OCR-VLM v1
ocrvlm_v1_q35
5 Oct 2026 — OCR page labels 7.33 $9.53 Winner 0.0096
OCR-VLM pilot
ocrvlm_pilot_q35
4 Oct 2026 — OCR page labels 1.37 $1.80 Keep —
OCR pilot 100k Running
ocr_pilot100k
— — — — $5.81 — —

Cost “—” means no cloud bill (our own GPU at home) or a cost that was not recorded; “Free” is free cloud compute. GPU-h is GPU hours.