Awesome Data Analysis
Section: Data Sources & Datasets · A lightweight library to easily share and access datasets for audio, computer vision, and NLP.
Entry
Appears in 6 awesome lists
The largest hub of ready-to-use NLP datasets for ML models with fast, easy-to-use and efficient data manipulation tools.
Section: Data Sources & Datasets · A lightweight library to easily share and access datasets for audio, computer vision, and NLP.
Section: Official Libraries · The largest hub of ready-to-use NLP datasets for ML models with fast, easy-to-use and efficient data manipulation tools.
Section: Libraries · standardized loaders and processing for thousands of NLP datasets.
Section: 9. Evaluation, Benchmarks & Datasets · Largest open repository of datasets.
Section: Other · 🤗 The largest hub of ready-to-use datasets for AI models with fast, easy-to-use and efficient data manipulation tools
Section: Datasets (35 · 22K) - The largest hub of ready-to-use datasets for AI models with fast,.. Apache-2 · (👨💻 710 · 🔀 3.3K · 📦 130K):
An awesome repository full of open datasets from an abundance of different categories.
(formerly known as pytorch-transformers and pytorch-pretrained-bert) provides state-of-the-art general-purpose architectures (BERT, GPT-2, RoBERTa, XLM, DistilBert, XLNet, CTRL...) for Natural Language Understanding (NLU) and Natural Language Generation (NLG) with over 32+ pretrained models in…
Open-source AI orchestration framework for building context-engineered, production-ready LLM applications. Design modular pipelines and agent workflows with explicit control over retrieval, routing, memory, and generation. Built for scalable agents, RAG, multimodal applications, semantic search,…
A curated list of resources dedicated to Natural Language Processing and text processing for Ruby.
PyTorch's official computer vision library with 50+ pre-trained model architectures including ResNet, EfficientNet, Vision Transformers (ViT), ConvNeXt, and more. The de facto standard model zoo for PyTorch computer vision. BSD-3-Clause licensed.
Industrial-strength natural language processing with 75+ languages, transformer pipelines, and production-grade NER, parsing, and text classification.
Hugging Face's tokenizers for modern NLP pipelines (original implementation) with bindings for Python.