Published on October 22, 2025, the post presents an OCR model for PDFs and scans, supporting tables, equations, and handwriting. Its approach combines synthetic data and unit tests as verifiable rewards.
On October 22, 2025, a post presented olmOCR 2, a model intended to convert PDFs and scans into clean text. According to the post, it supports tables, equations, and handwriting, and uses synthetic data and unit tests as verifiable rewards.
The report may interest teams evaluating OCR for complex documents and ways to check results. To confirm the scope and findings, consult the original post and review its technical details; the available text provides no performance metrics.