Back to News
Datalab•July 20, 2026
Marker
Featured on Blog
Open Source
OCR
Document Intelligence
PDF to Markdown
Machine Learning
Featured
🔥 This release made it to our blog
Datalab Launches Marker 2.0: An Open-Source PDF-to-Markdown Tool Built for RAG Pipeline
Datalab has released Marker 2.0, an open-source tool designed to convert PDFs to Markdown for RAG pipelines with increased speed and accuracy.
Read the Full Story 5 min read
### TL;DR
Marker is a high-performance document intelligence pipeline designed to convert PDFs, images, and various office documents into clean Markdown, JSON, HTML, or chunks. It leverages advanced deep learning models, including the Surya VLM, to handle complex layouts, tables, equations, and multi-language text extraction with high accuracy.
Key Insights & Metrics
Pricing
The code is licensed under Apache 2.0 (free for commercial use)
Cost structure
Version
2.0.0
Current release version
Hardware
Requires Python 3.10+ and PyTorch. For GPU acceleration, NVIDIA GPU with Docker and NVIDIA Container Toolkit is recommended. For CPU/Apple Silicon, llama.cpp is required for the inference backend.
Compute requirements
Category
Open Source
Licensing model
Region
United States
Primary region
Key Features
- Converts PDF, image, PPTX, DOCX, XLSX, HTML, and EPUB files
- Formats tables, equations, inline math, and code blocks
- Extracts and saves images while removing headers/footers
- Supports optional LLM-based accuracy boosting
- Works on GPU, CPU, or MPS hardware
- Multilingual support for global document processing
Discussion
0
Upvotes
0
Downvotes
0 reviews
Sign in to leave a review
Reviews
No reviews yet. Be the first to review!