On October 20, 2025, FinePdfs announced source code, labeled datasets, and a classifier intended for PDF OCR workflows.
Published on October 20, 2025, FinePdfs' announcement describes releasing the complete OCR-Annotations source code and a dataset of 1,600 labeled PDFs, as well as Gemma-LID-Annotation, with 20,000 samples per language. It also includes XGB-OCR, an OCR classifier for PDFs.
According to the announcement, the datasets and classifier may help engineers build or evaluate PDF OCR workflows. To check the details and verify the original material, consult FinePdfs' publication and inspect the cited repositories, data, and code; do not assume results beyond those reported. If you use AI to study or apply these resources, avoid sending documents containing personal data unless necessary, and check your organization's policy.