Coming Soon

Unlimited-OCR

Next-generation OCR engine by Baidu. One-shot long-horizon document parsing with a 32K context window — covering single images, multi-page documents, and PDFs. Coming soon to PDF to MD.

One-shot Parsing32K ContextPDF & ImageMIT LicenseSGLang Serving
32,768 Token context length
300 DPI PDF render precision
2 Modes Gundam / Base modes
MIT Open-source license

What is Unlimited-OCR

Unlimited-OCR is an open-source OCR system by Baidu, positioned as a next-generation document parsing solution beyond DeepSeek-OCR. It processes ultra-long documents in a single inference pass, achieving true One-shot Long-horizon Parsing.

🎯 Single Image OCR

Two processing modes: Gundam (base_size=1024, with cropping) for complex layouts, and Base (image_size=1024) for standard documents.

📄 Multi-page Documents

Supports multi-image sequence input with automatic multi-page joint parsing. Anti-repeat mechanism avoids redundant output.

📑 Direct PDF Parsing

Renders PDF pages to 300 DPI images via PyMuPDF for multi-page parsing. No preprocessing needed.

⚡ High-Performance Inference

Dual backends: Hugging Face Transformers and SGLang. SGLang provides OpenAI-compatible API with FlashAttention 3.

🔧 Batch Processing

Built-in infer.py script auto-launches SGLang server with concurrent requests. Supports image directories or PDF input.

🏗️ Tech Stack

PyTorch with bfloat16 precision on NVIDIA GPU. Requires Python 3.12, CUDA 12.9, Transformers 4.57.1.

Coming to PDF to MD

We are integrating Unlimited-OCR into the PDF to MD platform, bringing more powerful document parsing capabilities. Once integrated, you can access all features through our API.

🌐 Online Experience

Upload documents directly on the PDF to MD website and select the Unlimited-OCR engine to experience one-shot long-document parsing.

🔌 Developer API

RESTful API coming soon — supporting single image, multi-page, and PDF calls with streaming, batch processing, and custom parameters.

📊 Structured Output

Parsed results output directly as Markdown, preserving tables, lists, and heading hierarchies for seamless LLM workflow integration.

🛡️ Stable & Reliable

Built on the SGLang high-performance inference framework with auto-scaling and fault recovery for enterprise-grade OCR service.

How It Works

From document upload to structured output — fully automated.

01 Upload PDF / Image
02 Unlimited-OCR Parsing
03 Structured Markdown
04 API / Download Output

Get Early Access

The Unlimited-OCR API is launching soon. If you would like early access, pricing details, or technical collaboration, please reach out.