Skip to content
EN

OCR and document AI

Reading PDFs, scans and tables with models, and how to check the result.

15 links, newest first.

Get the weekly briefing

The best new links of the topics you pick, summarized with the source. At most one email a week.

Topics: OCR and document AI

Before the first issue we email you to confirm; leaving takes one click. Sent with CommsHarbor. Privacy

  1. Open-source classifier for tax document pages

    The GitHub project is a tax document page classifier built on Jev decisions. Its description claims 100% strict accuracy across 261 IRS forms at about $0.001 per page.

    Engineers can inspect an open-source approach to classifying tax document pages and its reported accuracy and cost.

  2. WeVisDoc and jina-ocr-v1 are new OCR models

    The post reports Tencent’s WeVisDoc 2B and 4B models under Apache 2.0 and jinaai’s jina-ocr-v1 under a non-commercial license. It says WeVisDoc-4B scores 95.38 on OmniDocBench v1.6, while NaviDC-OCR is current SOTA.

    The model options, license terms, and benchmark result help engineers compare OCR candidates.

  3. iOS OCR server using Apple’s Vision framework

    The open-source iOS-OCR-Server project uses Apple’s Vision framework to provide OCR from an iPhone. The post describes local processing and an API for use by devices on the same network.

    It offers engineers a way to explore on-device OCR through a network-accessible API.

  4. Testing Unlimited OCR Against an OCR Benchmark

    The author says Baidu Unlimited OCR was added to mlx-vlm and plans to test it against an OCR benchmark. The post also asks for other strong OCR benchmarks.

    Engineers evaluating OCR models can use the linked benchmark as a starting point.

  5. Unlimited OCR for Long Documents

    Baidu’s open-source OCR model uses Reference Sliding Window Attention to process documents longer than 40 pages in one forward pass. The post reports 3B total parameters, 500M activated, and results on OmniDocBench v1.5 and v1.6.

    Engineers can inspect the implementation of an approach to long-document OCR with constant KV Cache size.

  6. PaddleOCR 3.7 adds ONNX Runtime support

    PaddleOCR 3.7 adds an ONNX Runtime inference backend and support for CPU inference via OpenVINO and GPU inference via CUDA and TensorRT. The post says PP-OCRv6 models are up to 3.9× faster on CPU.

    Engineers can compare inference backends and hardware options when deploying OCR across platforms.

  7. PaddleOCR releases PP-OCRv6 model series

    PP-OCRv6 includes Tiny, Small, and Medium models ranging from 1.5M to 34.5M parameters. PaddlePaddle reports support for 50 languages, new OCR scenarios, and up to 5.2× faster CPU inference with OpenVINO.

    Engineers can compare model sizes, language coverage, and inference options for OCR deployments.

  8. Baidu releases PP-OCRv6 for multilingual OCR

    The post says PP-OCRv6 comes in 1.5M, 7.7M, and 34.5M parameter sizes, supports 48+ languages and several text types, and is designed for edge deployment.

    The model-size options and stated text coverage may help engineers assess it for OCR workloads.

  9. PaddleOCR-VL 1.6 reports 96.33% on OmniDocBench v1.6

    The post announces PaddleOCR-VL 1.6 and reports a 96.33% score on OmniDocBench v1.6, which it describes as a new state of the art.

    The benchmark result may help engineers track OCR model performance on document understanding.

  10. GLM-OCR: a 0.9B multimodal OCR model

    The post announces GLM-OCR, described as a 0.9B model using a multimodal GLM-V architecture and an MIT license. The author claims it ranks first on OmniDocBench v1.5 with a score of 94.62.

    Engineers evaluating OCR models can inspect its model page and compare the reported benchmark claim.

  11. Ente’s on-device OCR for Android

    Ente describes how it brought private, on-device text recognition to Ente Photos. The post says its OCR is fully open source.

    Useful as an implementation example for engineers exploring private, on-device OCR.

  12. olmOCR 2 Uses Verifiable Rewards for Document OCR

    The post introduces olmOCR 2, a model for converting PDFs and scans into clean text, with support for tables, equations, and handwriting. It says the model uses synthetic data and unit tests as verifiable rewards.

    The approach may interest engineers evaluating OCR for challenging documents and ways to verify model output.

  13. FinePdfs releases OCR datasets, models, and source code

    FinePdfs is releasing its full source code, OCR-Annotations with 1.6k labeled PDFs, Gemma-LID-Annotation with 20k samples per language, and XGB-OCR, an OCR classifier for PDFs.

    The datasets and classifier may help engineers build or evaluate PDF OCR workflows.

  14. Zerox extracts structured data from documents with vision models

    The post describes Zerox as an open-source OCR and document extraction project. It supports extracting specific information into a structured format using a schema, rather than converting the full document to Markdown.

    Schema-based extraction can help engineers retrieve targeted fields from documents.

  15. olmOCR extracts plain text from PDFs

    AllenAI introduces olmOCR, an open-source tool for extracting plain text from PDFs. The post says it handles many document types and can run on a GPU.

    Engineers evaluating PDF text-extraction pipelines can investigate an open-source alternative.

Build with AgentLog

List your MCP, skill or plugin

Reach the engineers who read these briefings.

Sponsor AgentLog

Footer, sidebar or featured slot for 30 days.

From US$ 60

See the slots

Send your own newsletter

CommsHarbor keeps contacts, consent and one-click unsubscribe together.

Free workspace

Open CommsHarbor