Title of a publication / dataset / resource title
, [code/data/Website ]
A curated list of resources for Document Understanding (DU) topic
This page lists names, links and short descriptions. The original list on GitHub is the source and belongs to its authors.
, [code/data/Website ]
, [Website] [benchmark] [code ]
dataset consists of 400,000 grayscale images in 16 classes, with 25,000 images per class
a portal to millions of documents created by industries that influence public health, hosted by the UCSF Library
from the Intelligent Sensory Information Systems, University of Amsterdam
dataset consists of documents from the states' lawsuit against the tobacco industry in the 1990s, consists of around 7 million documents
is a pure python library to read, write and manipulate PDF documents. It represents a PDF document as a JSON-like datastructure of nested lists, dictionaries and primitives (numbers, string, booleans, etc).
PDF Annotations with Labels and Structure is software that makes it easy to collect a series of annotations associated with a PDF document
Plumb a PDF for detailed information about each text character, rectangle, and line. Plus: Table extraction and visual debugging
Pdfminer.six is a community maintained fork of the original PDFMiner. It is a tool for extracting information from PDF documents. It focuses on getting and analyzing text data
Layout Parser is a deep learning based tool for document image layout analysis tasks
Table extraction from images
OCRmyPDF adds an OCR text layer to scanned PDF files, allowing them to be searched or copy-pasted
The Apache PDFBox library is an open source Java tool for working with PDF documents. This project allows creation of new PDF documents, manipulation of existing documents and the ability to extract content from documents
This project allows users to read and extract text and other content from PDF files. In addition the library can be used to create simple PDF documents containing text and geometrical shapes. This project aims to port PDFBox to C#
Resources and worksheet for the NICAR 2016 workshop of the same name
PDF tools benchmark
checking if pdf is born-digital
Apache2-licensed, PDF annotating platform for visually-rich documents that preserves the original layout and exports x,y positional data for tokens as well as span starts and stops. Based on PAWLs, but with a Python-based backend and readily deployable on your local machine, company intranet or…
deepdoctection is a Python library that orchestrates document extraction and document layout analysis tasks for images and pdf documents using deep learning models. It does not implement models but enables you to build pipelines using highly acknowledged libraries for object detection, OCR and…
Pydoxtools is an AI-composition library for dpocument analysis. It features an extensive toolset for building complex document analysis pipelines and recognizes most document formats out of the box. It supports typical NLP tasks such as keywords, summarization, question_answering out of the box.…
, 2021
The Microsoft Azure Machine Learning platform provides capabilities such as natural language processing, recommendation engine, pattern recognition, computer vision, and predictive modeling.
team-first hosted and on-prem text, image and PDF annotation tool powered by active learning, freemium based, costs $
This list contains links to great software tools and libraries and literature related to Optical Character Recognition (OCR).
Information retrieval resources
A curated list of awesome open source libraries to deploy, monitor, version and scale your machine learning
Curated resources and tools for applied machine learning in industry.
An awesome repository full of open datasets from an abundance of different categories.
A ranked list of awesome Python libraries for natural language processing (NLP).
BERT (Transformer, transfer learning) has catalyzed research in pretrained language models (PLMs) and has sparked many extensions. This repo contains a list of papers on PLMs.
컴퓨터 비전 자료모음 (Awesome 계열)
“The practice of software engineering, and its history is, itself, a complex study in humanity, coordination, and communication.”
hesreallyhim/awesome-claude-code
A hand-picked collection of the finest of resources for the most awesome of agents, Claude Code, the undisputed champion of coding companions, from the unstoppable team…
VoltAgent/awesome-agent-skills
A curated collection of 1000+ agent skills from official dev teams and the community, compatible with Claude Code, Codex, Gemini CLI, Cursor, and more.
josephmisiti/awesome-machine-learning
A curated list of awesome Machine Learning frameworks, libraries and software.
EthicalML/awesome-production-machine-learning
A curated list of awesome open source libraries to deploy, monitor, version and scale your machine learning
academic/awesome-datascience
:memo: An awesome Data Science repository to learn and apply for real world problems.
analysis-tools-dev/static-analysis
⚙️ A curated list of static analysis (SAST) tools and linters for all programming languages, config files, build tools, and more. The focus is on tools which improve…