Back to News
DatalabJuly 20, 2026

Marker

Featured on Blog
Open Source
OCR
Document Intelligence
PDF to Markdown
Machine Learning

Explore Marker

Visit the official website to learn more and get started

Featured
Datalab Launches Marker 2.0: An Open-Source PDF-to-Markdown Tool Built for RAG Pipeline

🔥 This release made it to our blog

Datalab Launches Marker 2.0: An Open-Source PDF-to-Markdown Tool Built for RAG Pipeline

Datalab has released Marker 2.0, an open-source tool designed to convert PDFs to Markdown for RAG pipelines with increased speed and accuracy.

### TL;DR

Marker is a high-performance document intelligence pipeline designed to convert PDFs, images, and various office documents into clean Markdown, JSON, HTML, or chunks. It leverages advanced deep learning models, including the Surya VLM, to handle complex layouts, tables, equations, and multi-language text extraction with high accuracy.

Key Insights & Metrics

Pricing
The code is licensed under Apache 2.0 (free for commercial use)
Cost structure
Version
2.0.0
Current release version
Hardware
Requires Python 3.10+ and PyTorch. For GPU acceleration, NVIDIA GPU with Docker and NVIDIA Container Toolkit is recommended. For CPU/Apple Silicon, llama.cpp is required for the inference backend.
Compute requirements
Category
Open Source
Licensing model
Region
United States
Primary region

Key Features

  • Converts PDF, image, PPTX, DOCX, XLSX, HTML, and EPUB files
  • Formats tables, equations, inline math, and code blocks
  • Extracts and saves images while removing headers/footers
  • Supports optional LLM-based accuracy boosting
  • Works on GPU, CPU, or MPS hardware
  • Multilingual support for global document processing

Discussion

0
Upvotes
0
Downvotes
0 reviews

Sign in to leave a review

Reviews

No reviews yet. Be the first to review!

🚀 Join the AI dev community — follow us everywhere

© 2026 MARKTECHPOST AI MEDIA INC. All rights reserved.Terms & ConditionsPrivacy Policy
Beta Mode