Skip to content

Entry

opendatalab/MinerU

Appears in 5 awesome lists

High-accuracy document parsing for LLM and RAG workflows. Converts PDFs, Word, PPTs, and images into structured Markdown/JSON with VLM+OCR dual engine.

Open github.comopendatalab/mineru

Found in these lists

Awesome Ai For Science

Section: High-Performance Document Processing · SOTA multimodal document parsing with 1.2B parameters outperforming GPT-4o, converts PDFs to LLM-ready Markdown/JSON

FreshScore 86

Awesome Generative AI

Section: Everything to Markdown to LLMs · A high-quality tool for convert PDF to Markdown and JSON

SlowScore 62

Awesome LLM Resources

Section: 数据 Data · MinerU is a one-stop, open-source, high-quality data extraction tool, supports PDF/webpage/e-book extraction.

FreshScore 87

Awesome Open Source AI

Section: 5. Retrieval-Augmented Generation (RAG) & Knowledge · High-accuracy document parsing for LLM and RAG workflows. Converts PDFs, Word, PPTs, and images into structured Markdown/JSON with VLM+OCR dual engine.

FreshScore 89

awesome-python

Section: Text Processing · Transforms complex documents like PDFs and Office docs into LLM-ready markdown/JSON for your Agentic workflows.

FreshScore 81

MarkItDown

Python tool for converting files and office documents to Markdown. Supports PDF, PowerPoint, Word, Excel, images, audio, HTML, and more with OCR and transcription capabilities. MIT licensed.

In 9 listsDetails

Crawl4AI

Open-source web crawler and scraper for LLMs and AI agents: any website into clean, LLM-ready Markdown. Run it yourself, or use Crawl4AI Cloud with one key.

In 5 listsDetails

Docling

Document processing toolkit for turning PDFs and other files into structured data for GenAI workflows.

In 5 listsDetails

olmOCR

Toolkit for linearizing academic PDFs into LLM-ready text with high accuracy and structure preservation, optimized for scientific literature extraction

In 5 listsDetails

PaddleOCR

Turn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/PDFs and LLMs. Supports 100+ languages.

In 5 listsDetails

TensorZero

(label: good-first-issue) TensorZero creates a feedback loop for optimizing LLM applications — turning production data into smarter, faster, and cheaper models.

In 4 listsDetails

Gitingest

Turn any Git repository into a simple text digest of its codebase. This is useful for feeding a codebase into any LLM.

In 4 listsDetails

Marker

Fast, accurate PDF-to-markdown converter with table extraction, equation handling, and optional LLM enhancement for RAG pipelines.

In 4 listsDetails