Long-document OCR: model processes more than 40 pages
Published on June 23, 2026, the report describes an open-source OCR model that processes documents over 40 pages in one pass and reports results on OmniDocBench v1.5 and v1.6.
Published on June 23, 2026, the report describes an open-source OCR model that uses Reference Sliding Window Attention to process documents over 40 pages in a single pass. According to the post, the model has 3 billion parameters in total, with 500 million activated, and reports results on OmniDocBench v1.5 and v1.6. The approach targets long-document OCR with a constant-size KV Cache.
To verify the details, consult the original post and inspect the implementation, method description, and reported benchmark results. If you use AI to study or apply this approach with organizational documents, follow your internal data policy: Rota Nacional detects personal data and applies the configured policy before model execution; usage metadata stays in the dashboard, while raw prompts and responses are kept in the encrypted audit archive.