Skip to content
92

AWESOME DATA SCIENCE

:memo: An awesome Data Science repository to learn and apply for real world problems.

30k stars6,658 forks881 entriesLast push Sep 30, 2026 (today)License MIT

This page lists names, links and short descriptions. The original list on GitHub is the source and belongs to its authors.

What is Data Science?

Data Science For Beginners

Microsoft are pleased to offer a 10-week, 20-lesson curriculum all about Data Science.

In 4 listsDetails

What is Data Science @ O'reilly

Data scientists combine entrepreneurship with patience, the willingness to build data products incrementally, the ability to explore, and the ability to iterate over a solution. They are inherently interdisciplinary. They can tackle all aspects of a problem, from initial data collection and data…

What is Data Science @ Quora

Data Science is a combination of a number of aspects of Data such as Technology, Algorithm development, and data interference to study the data, analyse it, and find innovative solutions to difficult problems. Basically Data Science is all about Analysing data and driving for business growth by…

The sexiest job of 21st century

Data scientists today are akin to Wall Street “quants” of the 1980s and 1990s. In those days people with backgrounds in physics and math streamed to investment banks and hedge funds, where they could devise entirely new algorithms and data strategies. Then a variety of universities developed…

Wikipedia

Data science is an interdisciplinary field that uses scientific methods, processes, algorithms and systems to extract knowledge and insights from many structural and unstructured data. Data science is related to data mining, machine learning and big data.

How to Become a Data Scientist

Data scientists are big data wranglers, gathering and analyzing large sets of structured and unstructured data. A data scientist’s role combines computer science, statistics, and mathematics. They analyze, process, and model data then interpret the results to create actionable plans for companies…

a very short history of #datascience

The story of how data scientists became sexy is mostly the story of the coupling of the mature discipline of statistics with a very young one--computer science. The term “Data Science” has emerged only recently to specifically designate a new profession that is expected to make sense of the vast…

Software Development Resources for Data Scientists

Data scientists concentrate on making sense of data through exploratory analysis, statistics, and models. Software developers apply a separate set of knowledge with different tools. Although their focus may seem unrelated, data science teams can benefit from adopting software development best…

Data Scientist Roadmap

Data science is an excellent career choice in today’s data-driven world where approx 328.77 million terabytes of data are generated daily. And this number is only increasing day by day, which in turn increases the demand for skilled data scientists who can utilize this data to drive business growth.

Navigating Your Path to Becoming a Data Scientist

_Data science is one of the most in-demand careers today. With businesses increasingly relying on data to make decisions, the need for skilled data scientists has grown rapidly. Whether it’s tech companies, healthcare organizations, or even government institutions, data scientists play a crucial…

Where do I Start?

Python

is by far the most popular language in science, due in no small part to the ease at which it can be used and the vibrant ecosystem of user-generated packages. To install packages, there are two main methods: Pip (invoked as pip install), the package manager that comes bundled with Python, and…

In 8 listsDetails

Agents >Frameworks

ADK-Rust

Production-ready AI agent development kit for Rust with model-agnostic design (Gemini, OpenAI, Anthropic), multiple agent types (LLM, Graph, Workflow), MCP support, and built-in telemetry.

Lumen

Agent framework for chatting with data, turning natural language into SQL, transformation pipelines and visualizations. Outputs are declarative specs that can be inspected, edited, reopened in a notebook or composed into a dashboard.

Agents >Tools

Frostbyte MCP

MCP server providing 13 data tools for AI agents: real-time crypto prices, IP geolocation, DNS lookups, web scraping to markdown, code execution, and screenshots. One API key for 40+ services.

Arch Tools

61 production-ready AI API tools for data science workflows: code analysis, web scraping, NLP, image generation, crypto data, and search. REST API and MCP protocol support. GitHub

Not Human Search

Search engine for AI agents that indexes 9,000+ AI tools and APIs, scoring each on agentic readiness (llms.txt, OpenAPI, MCP, ai-plugin.json). REST API and MCP server for programmatic tool discovery. GitHub

In 4 listsDetails

DeepAlpha

AI crypto trading framework using LightGBM + XGBoost ensemble with 72 ML features. 70.9% walk-forward validated accuracy on out-of-sample data. Supports Bybit and Binance. MIT licensed, available on PyPI.

In 3 lists

CAJAL

Local AI agent for generating publication-ready scientific papers with real arXiv citations, IMRaD structure, and tribunal scoring. Runs 100% offline via Ollama with 4B-9B models. MIT licensed. HuggingFace

ai-evaluation

Open-source LLM and agent evaluation framework with 50+ metrics, LLM-as-Judge augmentation, and guardrail scanners (jailbreak, PII, prompt-injection). Useful for scoring RAG outputs, agent trajectories, and function-calling behavior in data-science workflows.

In 2 lists

Kitaru

Open-source platform that records real AI agent runs, replays them against changes, and evaluates outcomes before deployment.

In 2 lists

Jev Social

Read-only social research agent that lets Jev choose bounded Instagram, TikTok, and LinkedIn operations, runs them through the local socai CLI in Chrome, and preserves source-linked evidence beside a cited report.

In 8 listsDetails

YYLO Benchmark

Open-source trusted-host experiment runner for historical agent tasks, supplied coding prompts, and workflows. Runs independent attempts with retained outputs, and compares models, harnesses, and configurations with different checks or judges later. MIT licensed.

In 2 lists

YYLO

Open-source command-line orchestrator for coding agents and repeatable workflows, with typed task, validation, merge, and release-readiness boundaries and receipt-backed repository changes. MIT licensed, installable via npm.

In 6 listsDetails

Agents >Research & Knowledge Retrieval

BGPT MCP

MCP server that gives AI agents access to a database of scientific papers built from raw experimental data extracted from full-text studies. Returns 25+ structured fields per paper including methods, results, sample sizes, and quality scores. GitHub

In 2 lists

Chunk Tuner

Open-source Python library and MCP server to benchmark document chunking strategies for RAG, score retrieval quality, and recommend configurations for a corpus.

II-Commons

Daily-updated skill and CLI for deterministic retrieval across arXiv, PubMed/PMC, and supported US policy corpora.

In 3 lists

Spraay x402 Gateway

x402 payment gateway with 23 Research & Reference endpoints for AI agents: Wikipedia, arXiv, PubMed, Wikidata, academic citation lookup, entity extraction, and more. Pay-per-call in USDC on Base & Solana — no API keys or subscriptions. Also serves 150+ endpoints across 39 categories including…

Suppr

AI literature search, document translation, and deep-research workspace for researchers.

Agents >Workflow

sim

Sim Studio's interface is a lightweight, intuitive way to quickly build and deploy LLMs that connect with your favorite tools.

Projects

Synthetic Hospital

A Medical Benchmark & EHR Simulation Platform

Training Resources >Tutorials

1000 Data Science Projects

you can run on the browser with IPython.

#tidytuesday

A weekly data project aimed at the R ecosystem.

In 3 lists

Data science your way

DataCamp Cheatsheets

Cheatsheets for data science.

PySpark Cheatsheet

Common PySpark patterns.

In 2 lists

Machine Learning, Data Science and Deep Learning with Python

LiveVideo tutorial that covers machine learning, Tensorflow, artificial intelligence, and neural networks.

In 3 lists

TutorialSearch

Free cross-platform search engine indexing 50,000+ tutorials from Udemy, Skillshare, Pluralsight, and other major learning platforms across 45+ categories.

In 2 lists

Your Guide to Latent Dirichlet Allocation

In 2 lists

Tutorials of source code from the book Genetic Algorithms with Python by Clinton Sheppard

Tutorials to get started on signal processing for machine learning

Realtime deployment

Tutorial on Python time-series model deployment.

Python for Data Science: A Beginner’s Guide

Minimum Viable Study Plan for Machine Learning Interviews

Understand and Know Machine Learning Engineering by Building Solid Projects

12 free Data Science projects to practice Python and Pandas

Best CV/Resume for Data Science Freshers

A collection of resume examples and tips tailored for data scientists.

In 2 lists

Understand Data Science Course in Java

Data Analytics Interview Questions (Beginner to Advanced)

In 2 lists

Top 100+ Data Science Interview Questions and Answers

DataDriven - SQL, Python, and Data Modeling Interview Questions

Interview practice with SQL query execution, Python, and data modeling exercises.

In 2 lists

StepByStepML

Interactive calculator that visualizes the step-by-step manual math behind machine learning algorithms for exam prep.

How to Build Optimal AI Agents That Actually Work

A developer handbook on designing and building effective AI agents.

Train LLM From Scratch

A straightforward method for training your LLM, from downloading data to generating text.

In 2 lists

Training Resources >Free Courses

Data Science

Open Source Society University

In 3 lists

Data Scientist with R

Data Scientist with Python

Genetic Algorithms OCW Course

AI Expert Roadmap

Roadmap to becoming an Artificial Intelligence Expert

In 4 listsDetails

Convex Optimization

Convex Optimization (basics of convex analysis; least-squares, linear and quadratic programs, semidefinite programming, minimax, extremal volume, and other problems; optimality conditions, duality theory...)

Learning from Data

Introduction to machine learning covering basic theory, algorithms and applications

Kaggle

Learn about Data Science, Machine Learning, Python etc

In 2 lists

ML Observability Fundamentals

Learn how to monitor and root-cause production ML issues.

Weights & Biases Effective MLOps: Model Development

Free Course and Certification for building an end-to-end machine using W&B

Python for Data Science by Scaler

This course is designed to empower beginners with the essential skills to excel in today's data-driven world. The comprehensive curriculum will give you a solid foundation in statistics, programming, data visualization, and machine learning.

MLSys-NYU-2022

Slides, scripts and materials for the Machine Learning in Finance course at NYU Tandon, 2022.

In 2 lists

Hands-on Train and Deploy ML

A hands-on course to train and deploy a serverless API that predicts crypto prices.

In 2 lists

LLMOps: Building Real-World Applications With Large Language Models

Learn to build modern software with LLMs using the newest tools and techniques in the field.

In 2 lists

Prompt Engineering for Vision Models

Learn to prompt cutting-edge computer vision models with natural language, coordinate points, bounding boxes, segmentation masks, and even other images in this free course from DeepLearning.AI.

In 3 lists

Data Science Course By IBM

Free resources and learn what data science is and how it’s used in different industries.

Neural Networks: Zero to Hero

A free video series by Andrej Karpathy covering neural networks from scratch — backpropagation, makemore, GPT, and more.

In 3 lists

Training Resources >MOOC's

Coursera Introduction to Data Science

Data Science - 9 Steps Courses, A Specialization on Coursera

43 weeks

In 2 lists

Data Mining - 5 Steps Courses, A Specialization on Coursera

30 weeks

In 2 lists

Machine Learning – 5 Steps Courses, A Specialization on Coursera

CS 109 Data Science

OpenIntro

CS 171 Visualization

Process Mining: Data science in Action

Oxford Deep Learning

Oxford Deep Learning - video

, Video Lectures Summer School Montreal

In 3 lists

Oxford Machine Learning

UBC Machine Learning - video

Data Science Specialization

🆓

In 2 lists

Coursera Big Data Specialization

30 weeks

In 2 lists

Statistical Thinking for Data Science and Analytics by Edx

Cognitive Class AI by IBM

Udacity - Deep Learning

In 2 lists

Keras in Motion

Microsoft Professional Program for Data Science

COMP3222/COMP6246 - Machine Learning Technologies

CS 231 - Convolutional Neural Networks for Visual Recognition

Fei-Fei Li, Andrej Karphaty, Justin Johnson (Stanford University)

In 2 lists

Coursera Tensorflow in practice

In 2 lists

Coursera Deep Learning Specialization

New series of 5 Deep Learning courses by Andrew Ng, now with Python rather than Matlab/Octave, and which leads to a specialization certificate.

In 6 listsDetails

365 Data Science Course

Coursera Natural Language Processing Specialization

Teaches how to use machine learning to understand and manipulate human language. It requires a working knowledge of machine learning, intermediate Python experience including DL frameworks & proficiency in calculus, linear algebra, & statistics.

In 3 lists

Coursera GAN Specialization

Codecademy's Data Science

Interactive courses for learning data analysis.

In 2 lists

Linear Algebra

Linear Algebra course by Gilbert Strang

A 2020 Vision of Linear Algebra (G. Strang)

Python for Data Science Foundation Course

Data Science: Statistics & Machine Learning

Machine Learning Engineering for Production (MLOps)

In 2 lists

Recommender Systems Specialization from University of Minnesota

is an intermediate/advanced level specialization focused on Recommender System on the Coursera platform.

Stanford Artificial Intelligence Professional Program

Data Scientist with Python

Programming with Julia

Scaler Data Science & Machine Learning Program

Data Science Skill Tree

Data Science for Beginners - Learn with AI tutor

Machine Learning for Beginners - Learn with AI tutor

Introduction to Data Science

Getting Started with Python for Data Science

Google Advanced Data Analytics Certificate

Professional courses in data analysis, statistics, and machine learning fundamentals.

Maschinelle Sprachgebrauchsanalyse - Grundlagen der Korpuslinguistik

course material on text-mining / corpus-linguistics in German funded by the federal state of North Rhine-Westphalia

Programmieren für Germanist*innen

course material: programming in python in German for digital humanities - funded by the federal state of North Rhine-Westphalia

QuiddityML

Short lessons with hands-on coding exercises and spaced repetition, covering Python, PyTorch, math for ML, ML foundations, NLP, and computer vision.

In 3 lists

Training Resources >Intensive Programs

Great Learning Data Science Programs

A collection of online data science and analytics certificate, postgraduate, and degree programs.

S2DS

WorldQuant University Applied Data Science Lab

Training Resources >Colleges

A list of colleges and universities offering degrees in data science.

Data Science Degree @ Berkeley

Data Science Degree @ UVA

Data Science Degree @ Wisconsin

BS in Data Science & Applications

MS in Computer Information Systems @ Boston University

MS in Business Analytics @ ASU Online

MS in Applied Data Science @ Syracuse

M.S. Management & Data Science @ Leuphana

Master of Data Science @ Melbourne University

Msc in Data Science @ The University of Edinburgh

Master of Management Analytics @ Queen's University

Master of Data Science @ Illinois Institute of Technology

Master of Applied Data Science @ The University of Michigan

Master Data Science and Artificial Intelligence @ Eindhoven University of Technology

Master's Degree in Data Science and Computer Engineering @ University of Granada

The Data Science Toolbox >Comparison

datacompy

DataComPy is a package to compare two Pandas DataFrames.

In 4 listsDetails

Regression

Linear Regression

is an approach to model the relationship between two variables by fitting a linear equation to observed data. One variable is considered to be an explanatory variable, and the other is considered to be a dependent variable.

In 3 lists

Ordinary Least Squares

Logistic Regression

Easily plottable and understandable classification.

In 3 lists

Stepwise Regression

Multivariate Adaptive Regression Splines

Softmax Regression

Locally Estimated Scatterplot Smoothing

k-nearest neighbor

The prototypical clustering method.

In 2 lists

Support Vector Machines

Decision Trees

The tree provides an interpretation.

In 3 lists

ID3 algorithm

C4.5 algorithm

Ensemble Learning

Boosting

In 2 lists

Stacking

Bagging

Random Forest

AdaBoost

, Python Code

In 2 lists

Clustering

scikit-learn clustering algorithms.

In 2 lists

Fuzzy clustering

Mixture models

Dimension Reduction

Principal Component Analysis (PCA)

t-SNE; t-distributed Stochastic Neighbor Embedding

Neural Networks

Self-organizing map

Adaptive resonance theory

Hidden Markov Models (HMM)

Clustering

Q Learning

SARSA (State-Action-Reward-State-Action) algorithm

Temporal difference learning

k-Means

Apriori

EM (Expectation-Maximization)

PageRank

Naive Bayes

is a machine learning algorithm that is used solved calssification problems. It's based on applying Bayes' theorem with strong independence assumptions between the features.

In 4 listsDetails

CART (Classification and Regression Trees)

In 2 lists

XGBoost (Extreme Gradient Boosting)

LightGBM (Light Gradient Boosting Machine)

CatBoost

is a fast, scalable, high performance Gradient Boosting on Decision Trees library, used for ranking, classification, regression and other machine learning tasks for Python, R, Java, C++. Supports computation on CPU and GPU.

In 3 lists

HDBSCAN (Hierarchical Density-Based Spatial Clustering of Applications with Noise)

FP-Growth (Frequent Pattern Growth Algorithm)

Isolation Forest

Deep Embedded Clustering (DEC)

TPU (Top-k Periodic and High-Utility Patterns)

Context-Aware Rule Mining (Transformer-Based Framework)

Multilayer Perceptron

Convolutional Neural Network (CNN)

Recurrent Neural Network (RNN)

Boltzmann Machines

Autoencoder

Generative Adversarial Network (GAN)

Transformer

Conditional Random Field (CRF)

ML System Designs)

The Data Science Toolbox >General Machine Learning Packages

scikit-learn

A Python module for machine learning built on top of SciPy.

In 3 lists

scikit-multilearn

Multi-label classification for python.

In 3 lists

sklearn-expertsys

Highly interpretable classifiers for scikit learn.

In 2 lists

scikit-feature

Feature selection repository in Python.

In 2 lists

scikit-rebate

A scikit-learn-compatible Python implementation of ReBATE, a suite of Relief-based feature selection algorithms for Machine Learning.

In 2 lists

seqlearn

Sequence classification toolkit for Python.

In 2 lists

sklearn-bayes

Python package for Bayesian Machine Learning with scikit-learn API.

In 3 lists

sklearn-crfsuite

A scikit-learn-inspired API for CRFsuite.

In 2 lists

sklearn-deap

Use evolutionary algorithms instead of gridsearch in scikit-learn.

In 2 lists

sigopt_sklearn

sklearn-evaluation

Model evaluation made easy: plots, tables, and markdown reports.

In 2 lists

scikit-image

A collection of algorithms for image processing in Python.

In 6 listsDetails

scikit-opt

Swarm Intelligence in Python (Genetic Algorithm, Particle Swarm Optimization, Simulated Annealing, Ant Colony Algorithm, Immune Algorithm, Artificial Fish Swarm Algorithm in Python)

In 3 lists

scikit-posthocs

Post-hoc tests for statistical analysis of data.

In 3 lists

feature-engine

me_fasttext

Memory-efficient FastText variant with exact trie n-gram IDs, structure-aware row sharing, and mmap serving for large-vocabulary NLP.

pystruct

Simple structured learning framework for Python.

In 2 lists

Shogun

xLearn

A high performance, easy-to-use, and scalable machine learning package, which can be used to solve large-scale machine learning problems. xLearn is especially useful for solving machine learning problems on large-scale sparse data, which is very common in Internet services such as online…

In 3 lists

cuML

is a suite of libraries that implement machine learning algorithms and mathematical primitives functions that share compatible APIs with other RAPIDS projects. cuML enables data scientists, researchers, and software engineers to run traditional tabular ML tasks on GPUs without going into the…

In 7 listsDetails

causalml

Uplift modeling and causal inference with machine learning algorithms.

In 2 lists

mlpack

A scalable C++ machine learning library (Python bindings).

In 6 listsDetails

MLxtend

A library of extension and helper modules for Python's data analysis and machine learning libraries.

In 4 listsDetails

modAL

A modular active learning framework for Python, built on top of scikit-learn.

In 4 listsDetails

Sparkit-learn

PySpark + scikit-learn = Sparkit-learn.

In 2 lists

hyperlearn

50%+ Faster, 50%+ less RAM usage, GPU support re-written Sklearn, Statsmodels.

In 2 lists

dlib

zap: - A toolkit for making real world machine learning and data analysis applications in C++. [Boost] website

In 7 listsDetails

imodels

jSciPy

A Java port of SciPy's signal processing module, offering filters, transformations, and other scientific computing utilities.

In 3 lists

RuleFit

Implementation of the rulefit.

In 2 lists

pyGAM

A Python library for generalized additive models with built-in smoothing and regularization.

In 3 lists

Deepchecks

Validation & testing of machine learning models and data during model development, deployment, and production. This includes checks and suites related to various types of issues, such as model performance, data integrity, distribution mismatches, and more.

In 8 listsDetails

scikit-survival

interpretable

XGBoost

Scalable, Portable and Distributed Gradient Boosting (GBDT, GBRT or GBM) Library, for Python, R, Java, Scala, C++ and more. Runs on single machine, Hadoop, Spark, Flink and DataFlow. [Apache2]

In 11 listsDetails

LightGBM

Microsoft's fast, distributed, high performance gradient boosting (GBDT, GBRT, GBM or MART) framework based on decision tree algorithms, used for ranking, classification and many other machine learning tasks.

In 8 listsDetails

CatBoost

General purpose gradient boosting on decision trees library with categorical features support out of the box. It is easy to install, contains fast inference implementation and supports CPU and GPU (even multi-GPU) computation.

In 10 listsDetails

PerpetualBooster

A gradient boosting machine that doesn't need hyperparameter optimization, with a simple budget parameter to control model complexity.

In 3 lists

JAX

| Python | - Composable transformations of Python+NumPy programs: differentiate, vectorize, JIT to GPU/TPU, and more

In 6 listsDetails

PhilanthroPy

Scikit-learn native toolkit for nonprofit fundraising analytics: leakage-safe donor propensity, lapse, planned-giving, wealth-screening and revenue-forecasting estimators.

In 2 lists

The Data Science Toolbox >Deep Learning Packages

PyTorch

(label: good first issue) PyTorch is an open source machine learning library based on the Torch library, used for applications such as computer vision and natural language processing.

In 16 listsDetails

TorchDR

GPU and multi-GPU dimensionality reduction with a scikit-learn-compatible API.

In 2 lists

torchvision

PyTorch's official computer vision library with 50+ pre-trained model architectures including ResNet, EfficientNet, Vision Transformers (ViT), ConvNeXt, and more. The de facto standard model zoo for PyTorch computer vision. BSD-3-Clause licensed.

In 9 listsDetails

torchtext

:Torchtext是一个非常好用的库,可以帮助我们很好的解决文本的预处理问题。此github存储库包含两部分:; torchText.data:文本的通用数据加载器、抽象和迭代器(包括词汇和词向量); torchText.datasets:通用NLP数据集的预训练加载程序 我们只需要通过pip install torchtext安装好torchtext后,便可以开始体验Torchtext 的种种便捷之处。

In 4 listsDetails

torchaudio

A set of tools and building blocks for audio and speech processing, designed to accelerate the development and deployment of machine learning applications in these domains. It offers GPU-compatible, differentiable, and production-ready components, making it valuable for integrating audio…

In 6 listsDetails

ignite

High-level library for training and evaluating neural networks in PyTorch with an engine, events & handlers system for maximum flexibility. BSD-3-Clause licensed.

In 7 listsDetails

PyTorchNet

PyToune

A Keras-like framework for PyTorch that handles much of the boilerplating code needed to train neural networks.

In 2 lists

skorch

scikit-learn compatible neural network library that wraps PyTorch. Seamlessly integrate PyTorch models with scikit-learn pipelines, grid search, and cross-validation.

In 5 listsDetails

PyVarInf

Python package facilitating the use of Bayesian Deep Learning methods with Variational Inference for PyTorch.

In 3 lists

pytorch_geometric

Graph neural network library for PyTorch enabling molecular modeling, materials discovery, protein interaction networks, and scientific knowledge graph learning (23.7k+ stars)

In 7 listsDetails

GPyTorch

A highly efficient and modular implementation of Gaussian Processes in PyTorch.

In 3 lists

pyro

Deep universal probabilistic programming with Python and PyTorch.

In 3 lists

Catalyst

High-level utils for PyTorch DL & RL research. It was developed with a focus on reproducibility, fast experimentation and code/ideas reusing. Being able to research/develop something new, rather than write another regular train loop.

In 6 listsDetails

pytorch_tabular

Yolov3

PyTorch implementation of YOLOv3, YOLOv3-SPP, and YOLOv3-tiny for real-time object detection with training, validation, inference, and multi-format export.

In 3 lists

Yolov5

Ultralytics YOLOv5 in PyTorch for object detection, instance segmentation, classification, training, and export.

In 3 lists

Yolov8

Ultralytics YOLO27, YOLO26, YOLO11, YOLOv8 — object detection, instance segmentation, semantic segmentation, image classification, pose estimation, object tracking

In 6 listsDetails

OpenLanguageModel

PyTorch-native library for building, training and teaching transformer language models, with architectures written as ordinary nn.Modules.

TensorFlow

How to use the Hexagon Delegate to speed up model inference on mobile and edge devices. Also see blog post Accelerating TensorFlow Lite on Qualcomm Hexagon DSPs.

In 23 listsDetails

TensorLayer

Deep learning and reinforcement learning library for researchers and engineers

In 3 lists

TFLearn

Deep learning library featuring a higher-level API for TensorFlow.

In 5 listsDetails

Sonnet

Sonnet is DeepMind's library built on top of TensorFlow for building complex neural networks.

In 6 listsDetails

tensorpack

TRFL

TensorFlow Reinforcement Learning.

In 2 lists

Polyaxon

MLOps Tools For Managing & Orchestrating The Machine Learning LifeCycle. Reproducible and scalable machine learning workflows on Kubernetes with experiment tracking, model management, and pipeline orchestration. Apache 2.0 licensed.

In 9 listsDetails

NeuPy

tfdeploy

Deploy TensorFlow graphs for fast evaluation and export to TensorFlow-less environments running numpy.

In 2 lists

tensorflow-upstream

TensorFlow ROCm port.

In 2 lists

TensorFlow Fold

Deep learning with dynamic computation graphs in TensorFlow.

In 2 lists

tensorlm

TensorLight

A high-level framework for TensorFlow.

In 2 lists

Mesh TensorFlow

Model Parallelism Made Easier.

In 2 lists

Ludwig

Low-code framework for building custom LLMs and deep neural networks. Declarative YAML configuration for training state-of-the-art models with PEFT/LoRA, 4-bit quantization, distributed training via Hugging Face Accelerate, and native Kubernetes support. Linux Foundation AI project. Apache 2.0…

In 6 listsDetails

TF-Agents

A reliable, scalable and easy to use TensorFlow library for contextual bandits and reinforcement learning.

In 4 listsDetails

TensorForce

An open-source deep reinforcement learning framework, with an emphasis on modularized flexible library design and straightforward usability for applications in research and practice.

In 2 lists

Keras

Neural Networks on top of tensorflow, examples. keras-contrib - Keras community contributions. keras-tuner - Hyperparameter tuning for Keras. hyperas - Keras + Hyperopt: Convenient hyperparameter optimization wrapper. elephas - Distributed Deep learning with Keras & Spark. tflearn - Neural…

In 8 listsDetails

keras-contrib

Keras community contributions.

In 2 lists

Hyperas

Keras + Hyperopt: A very simple wrapper for convenient hyperparameter optimization.

In 4 listsDetails

Elephas

Distributed Deep learning with Keras & Spark.

In 3 lists

Hera

Spektral

Deep learning on graphs.

In 2 lists

qkeras

A quantization deep learning library.

In 2 lists

keras-rl

Deep Reinforcement Learning for Keras.

In 2 lists

Talos

Hyperparameter Optimization for TensorFlow, Keras and PyTorch.

In 3 lists

altair

Declarative statistical visualization library for Python. Can easily do many data transformation within the code to create graph

In 3 lists

amcharts

Three libraries for traditional charts, stock, and maps. Features a hand-drawn style theme option.

In 3 lists

anychart

AnyChart Component for Ember CLI provides an easy way to use AnyChart JavaScript Charts with Ember Framework

In 3 lists

bokeh

Python library for interactive data visualization in the browser, with support for networks.

In 2 lists

Comet

slemma

In 2 lists

cartodb

Cube

d3plus

D3's simpler, easier to use cousin. Mostly predefined templates that you can just plug data in.

In 2 lists

Data-Driven Documents(D3js)

Allows the user to manipulate documents based on data to render charts in SVG.

In 9 listsDetails

dygraphs

Interactive line charts library that works with huge datasets.

In 3 lists

exhibit

In 2 lists

gephi

Leading visualization and exploration software for all kinds of graphs and networks.

In 6 listsDetails

ggplot2

Resource for plotting a wide range of data (useful for visualizing survey data). Additional Information: GNU GENERAL PUBLIC LICENSE.

In 4 listsDetails

Glue

Google Chart Gallery

Highcharts

A charting library written in pure JavaScript, offering an easy way of adding interactive charts to your web site or web application.

In 6 listsDetails

import.io

.

In 3 lists

Matplotlib

is a 2D plotting library for creating static, animated, and interactive visualizations in Python. Matplotlib produces publication-quality figures in a variety of hardcopy formats and interactive environments across platforms.

In 4 listsDetails

nvd3

Netron

Visualizer for deep learning and machine learning models (no Python code, but visualizes models from most Python Deep Learning frameworks).

In 15 listsDetails

Openrefine

plot.ly

Easy-to-use web service that allows for rapid creation of complex charts, from heatmaps to histograms. Upload data to create and style charts with Plotly's online spreadsheet. Fork others' plots.

In 4 listsDetails

raw

SVG Data Visualization Generator - sunburst, circular dendrogram or multiple convex hull, for example. with tutorials: https://rawgraphs.io/learning

In 5 listsDetails

Resseract Lite

Seaborn

A Python visualization library based on matplotlib. It provides a high-level interface for drawing attractive statistical graphics.

In 6 listsDetails

techanjs

Stock and financial charts.

In 2 lists

Timeline

Easy-to-make, beautiful timelines.

In 4 listsDetails

variancecharts

vida

vizzu

Library for animated data visualizations and data stories.

In 3 lists

Wrangler

r2d3

NetworkX

Python package for the creation, manipulation, and study of the structure, dynamics, and functions of complex networks.

In 2 lists

Redash

"Redash has support for querying multiple databases, including: Redshift, Google BigQuery, PostgreSQL, MySQL, Graphite, Presto, Google Spreadsheets, Cloudera Impala, Hive and custom scripts."

In 6 listsDetails

Metabase

Easy way for everyone in your company to ask questions and learn from data. (Source Code) AGPL-3.0 Java/Docker

In 8 listsDetails

C3

customizable library based on D3.js for easy chart drawing.

In 4 listsDetails

TensorWatch

Debugging and visualization tool for machine learning and data science. It extensively leverages Jupyter Notebook to show real-time visualizations of data in running processes such as machine learning training.

In 5 listsDetails

geomap

Dash

is a popular Python framework for building ML & data science web apps for Python, R, Julia, and Jupyter.

In 2 lists

MetaReview

Free online meta-analysis platform with 11 interactive D3.js statistical charts (forest plot, funnel plot, Galbraith, L'Abbé, Baujat, etc.), 5 effect size measures, AI literature screening, and publication-ready report export. github.com

torchvista

Interactive notebook-based tool to visualize the forward pass of any PyTorch model.

In 3 lists

FlexViz

Python library for interactive, cross-filtered dashboards that stay responsive on 100M+ rows by aggregating with Polars on the server.

In 4 listsDetails

The Data Science Toolbox >Miscellaneous Tools

The Data Science Lifecycle Process

The Data Science Lifecycle Process is a process for taking data science teams from Idea to Value repeatedly and sustainably. The process is documented in this repo

Data Science Lifecycle Template Repo

Template repository for data science lifecycle project

TabGAN

Synthetic tabular data generation using GANs, Diffusion Models, and LLMs with adversarial filtering and privacy metrics.

In 2 lists

RexMex

A general purpose recommender metrics library for fair evaluation.

In 2 lists

ChemicalX

A PyTorch based deep learning library for drug pair scoring.

In 3 lists

FileShot.io

Secure zero-knowledge encrypted file sharing (AES-256-GCM in-browser). No account required, MIT licensed, self-hostable, optional link expiry.

In 3 lists

CorpusExplorer

Software for corpus linguists and text/data mining enthusiasts. Build your own corpora in over 60 languages. Use over 50 tools/visualizations.

In 2 lists

PyTorch Geometric Temporal

Representation learning on dynamic graphs.

In 5 listsDetails

Little Ball of Fur

A graph sampling library for NetworkX with a Scikit-Learn like API.

In 5 listsDetails

Karate Club

An unsupervised machine learning extension library for NetworkX with a Scikit-Learn like API.

In 7 listsDetails

ML Workspace

All-in-one web-based IDE for machine learning and data science. The workspace is deployed as a Docker container and is preloaded with a variety of popular data science libraries (e.g., Tensorflow, PyTorch) and dev tools (e.g., Jupyter, VS Code)

In 10 listsDetails

xonsh shell

A Python-powered shell that enables integration, management and orchestration of data science libraries mostly written in Python, allowing you to build pipelines, code and command-based workflows. It can also be used as a kernel for Jupyter Notebook.

In 4 listsDetails

Neptune.ai

Community-friendly platform supporting data scientists in creating and sharing machine learning models. Neptune facilitates teamwork, infrastructure management, models comparison and reproducibility.

In 6 listsDetails

steppy

Lightweight, Python library for fast and reproducible machine learning experimentation. Introduces very simple interface that enables clean machine learning pipeline design.

In 2 lists

steppy-toolkit

Curated collection of the neural networks, transformers and models that make your machine learning work faster and more effective.

Datalab from Google

easily explore, visualize, analyze, and transform data using familiar languages, such as Python and SQL, interactively.

Hortonworks Sandbox

is a personal, portable Hadoop environment that comes with a dozen interactive Hadoop tutorials.

R

is a free software environment for statistical computing and graphics.

In 2 lists

Tidyverse

is an opinionated collection of R packages designed for data science. All packages share an underlying design philosophy, grammar, and data structures.

RStudio

IDE – powerful user interface for R. It’s free and open source, and works on Windows, Mac, and Linux.

In 7 listsDetails

Python - Pandas - Anaconda

Completely free enterprise-ready Python distribution for large-scale data processing, predictive analytics, and scientific computing

In 4 listsDetails

Pandas GUI

Pandas GUI

NuriStat

Free open-source SPSS alternative — menu-driven desktop statistics (t-tests, ANOVA, regression, survival analysis, ROC) with SPSS .sav import/export

Polars

Fast DataFrame library for Rust and Python, designed as a faster alternative to Pandas

In 10 listsDetails

CiteMe

free academic citation generator with a built-in reference checker that flags fabricated or hallucinated references. Searches 11+ scholarly databases (OpenAlex, PubMed, Semantic Scholar, CrossRef, SciELO), formats 40+ citation styles, and offers a public API. No sign-up; available in English,…

In 2 lists

Scikit-Learn

Machine Learning in Python

In 5 listsDetails

NumPy

NumPy is fundamental for scientific computing with Python. It supports large, multi-dimensional arrays and matrices and includes an assortment of high-level mathematical functions to operate on these arrays.

In 7 listsDetails

Vaex

Vaex is a Python library that allows you to visualize large datasets and calculate statistics at high speeds.

SciPy

SciPy works with NumPy arrays and provides efficient routines for numerical integration and optimization.

In 7 listsDetails

Data Science Toolbox

Coursera Course

Data Science Toolbox

Blog

Wolfram Data Science Platform

Take numerical, textual, image, GIS or other data and give it the Wolfram treatment, carrying out a full spectrum of data science analysis and visualization and automatically generate rich interactive reports—all powered by the revolutionary knowledge-based Wolfram Language.

Datadog

Solutions, code, and devops for high-scale data science.

In 8 listsDetails

variancecharts

Kite Development Kit

The Kite Software Development Kit (Apache License, Version 2.0), or Kite for short, is a set of libraries, tools, examples, and documentation focused on making it easier to build systems on top of the Hadoop ecosystem.

Domino Data Labs

Run, scale, share, and deploy your models — without any infrastructure or setup.

In 4 listsDetails

Apache Flink

A platform for efficient, distributed, general-purpose data processing.

In 6 listsDetails

Apache Hama

Apache Hama is an Apache Top-Level open source project, allowing you to do advanced analytics beyond MapReduce.

Weka

Weka is a collection of machine learning algorithms for data mining tasks.

Octave

GNU Octave is a high-level interpreted language, primarily intended for numerical computations.(Free Matlab)

In 6 listsDetails

Apache Spark

Lightning-fast cluster computing

In 7 listsDetails

Hydrosphere Mist

a service for exposing Apache Spark analytics jobs and machine learning models as realtime, batch or reactive web services.

In 3 lists

Data Mechanics

A data science and engineering platform making Apache Spark more developer-friendly and cost-effective.

In 2 lists

Caffe

Deep Learning Framework

Torch

A SCIENTIFIC COMPUTING FRAMEWORK FOR LUAJIT

Nervana's python based Deep Learning Framework

Intel® Nervana™ reference deep learning framework committed to best performance on all hardware.

In 4 listsDetails

Skale

High performance distributed data processing in NodeJS

Aerosolve

A machine learning package built for humans.

Intel framework

Intel® Deep Learning Framework

Datawrapper

An open source data visualization platform helping everyone to create simple, correct and embeddable charts. Also at github.com

In 2 lists

Tensor Flow

TensorFlow is an Open Source Software Library for Machine Intelligence

In 5 listsDetails

Natural Language Toolkit

An introductory yet powerful toolkit for natural language processing and classification

In 4 listsDetails

FunASR

Industrial-grade speech recognition toolkit supporting 50+ languages with built-in VAD, punctuation, speaker diarization, and emotion detection. OpenAI-compatible API server included.

In 9 listsDetails

Annotation Lab

Free End-to-End No-Code platform for text annotation and DL model training/tuning. Out-of-the-box support for Named Entity Recognition, Classification, Relation extraction and Assertion Status Spark NLP models. Unlimited support for users, teams, projects, documents.

In 2 lists

nlp-toolkit for node.js

This module covers some basic nlp principles and implementations. The main focus is performance. When we deal with sample or training data in nlp, we quickly run out of memory. Therefore every implementation in this module is written as stream to only hold that data in memory that is currently…

Julia

high-level, high-performance dynamic programming language for technical computing

In 5 listsDetails

IJulia

a Julia-language backend combined with the Jupyter interactive environment

Apache Zeppelin

Web-based notebook that enables data-driven, interactive data analytics and collaborative documents with SQL, Scala and more

In 4 listsDetails

Featuretools

An open source framework for automated feature engineering written in python

In 6 listsDetails

Optimus

Cleansing, pre-processing, feature engineering, exploratory data analysis and easy ML with PySpark backend.

Albumentations

А fast and framework agnostic image augmentation library that implements a diverse set of augmentation techniques. Supports classification, segmentation, and detection out of the box. Was used to win a number of Deep Learning competitions at Kaggle, Topcoder and those that were a part of the CVPR…

In 3 lists

DVC

An open-source data science version control system. It helps track, organize and make data science projects reproducible. In its very basic scenario it helps version control and share large data and model files.

In 8 listsDetails

Lambdo

is a workflow engine that significantly simplifies data analysis by combining in one analysis pipeline (i) feature engineering and machine learning (ii) model training and prediction (iii) table population and column evaluation.

In 2 lists

Feast

A feature store for the management, discovery, and access of machine learning features. Feast provides a consistent view of feature data for both model training and model serving.

In 6 listsDetails

Polyaxon

MLOps Tools For Managing & Orchestrating The Machine Learning LifeCycle. Reproducible and scalable machine learning workflows on Kubernetes with experiment tracking, model management, and pipeline orchestration. Apache 2.0 licensed.

In 9 listsDetails

UBIAI

Easy-to-use text annotation tool for teams with most comprehensive auto-annotation features. Supports NER, relations and document classification as well as OCR annotation for invoice labeling

In 3 lists

Trains

Auto-Magical Experiment Manager, Version Control & DevOps for AI

In 3 lists

Hopsworks

Open-source data-intensive machine learning platform with a feature store. Ingest and manage features for both online (MySQL Cluster) and offline (Apache Hive) access, train and serve models at scale.

In 5 listsDetails

MindsDB

MindsDB is an Explainable AutoML framework for developers. With MindsDB you can build, train and use state of the art ML models in as simple as one line of code.

In 11 listsDetails

Lightwood

A Pytorch based framework that breaks down machine learning problems into smaller blocks that can be glued together seamlessly with an objective to build predictive models with one line of code.

In 2 lists

AWS Data Wrangler

An open-source Python package that extends the power of Pandas library to AWS connecting DataFrames and AWS data related services (Amazon Redshift, AWS Glue, Amazon Athena, Amazon EMR, etc).

In 3 lists

Amazon Rekognition

AWS Rekognition is a service that lets developers working with Amazon Web Services add image analysis to their applications. Catalog assets, automate workflows, and extract meaning from your media and applications.

In 2 lists

Amazon Textract

Automatically extract printed text, handwriting, and data from any document.

In 2 lists

Amazon Lookout for Vision

Spot product defects using computer vision to automate quality inspection. Identify missing product components, vehicle and structure damage, and irregularities for comprehensive quality control.

Amazon CodeGuru

Automate code reviews and optimize application performance with ML-powered recommendations.

In 2 lists

CML

An open source toolkit for using continuous integration in data science projects. Automatically train and test models in production-like environments with GitHub Actions & GitLab CI, and autogenerate visual reports on pull/merge requests.

In 4 listsDetails

Dask

An open source Python library to painlessly transition your analytics code to distributed computing systems (Big Data)

In 3 lists

DuckDB

An in-process SQL OLAP database management system

In 8 listsDetails

Statsmodels

A Python-based inferential statistics, hypothesis testing and regression framework

In 3 lists

Gensim

An open-source library for topic modeling of natural language text

In 3 lists

spaCy

A performant natural language processing toolkit

In 6 listsDetails

Grid Studio

Grid studio is a web-based spreadsheet application with full integration of the Python programming language.

In 2 lists

Python Data Science Handbook

Python Data Science Handbook: full text in Jupyter Notebooks

In 4 listsDetails

Shapley

A data-driven framework to quantify the value of classifiers in a machine learning ensemble.

In 4 listsDetails

DAGsHub

A platform built on open source tools for data, model and pipeline management.

In 4 listsDetails

Deepnote

A new kind of data science notebook. Jupyter-compatible, with real-time collaboration and running in the cloud.

In 3 lists

Valohai

An MLOps platform that handles machine orchestration, automatic reproducibility and deployment.

In 2 lists

PyMC3

A Python Library for Probabalistic Programming (Bayesian Inference and Machine Learning)

In 2 lists

PyStan

Python interface to Stan (Bayesian inference and modeling)

hmmlearn

Unsupervised learning and inference of Hidden Markov Models

Chaos Genius

ML powered analytics engine for outlier/anomaly detection and root cause analysis

In 5 listsDetails

PySAD

Python library for anomaly detection on streaming data

In 2 lists

Nimblebox

A full-stack MLOps platform designed to help data scientists and machine learning practitioners around the world discover, create, and launch multi-cloud apps from their web browser.

Towhee

A Python library that helps you encode your unstructured data into embeddings.

In 2 lists

LineaPy

Ever been frustrated with cleaning up long, messy Jupyter notebooks? With LineaPy, an open source Python library, it takes as little as two lines of code to transform messy development code into production pipelines.

In 2 lists

envd

🏕️ machine learning development environment for data science and AI/ML engineering teams

In 5 listsDetails

Explore Data Science Libraries

A search engine 🔎 tool to discover & find a curated list of popular & new libraries, top authors, trending project kits, discussions, tutorials & learning resources

MLEM

🐶 Version and deploy your ML models following GitOps principles

In 4 listsDetails

MLflow

MLOps framework for managing ML models across their full lifecycle

In 6 listsDetails

cleanlab

Python library for data-centric AI and automatically detecting various issues in ML datasets

In 8 listsDetails

AutoGluon

AutoML to easily produce accurate predictions for image, text, tabular, time-series, and multi-modal data

In 5 listsDetails

Arize AI

Arize AI community tier observability tool for monitoring machine learning models in production and root-causing issues such as data quality and performance drift.

In 3 lists

Aureo.io

Aureo.io is a low-code platform that focuses on building artificial intelligence. It provides users with the capability to create pipelines, automations and integrate them with artificial intelligence models – all with their basic data.

ERD Lab

Free cloud based entity relationship diagram (ERD) tool made for developers.

In 5 listsDetails

Arize-Phoenix

MLOps in a notebook - uncover insights, surface problems, monitor, and fine tune your models.

Comet

An MLOps platform with experiment tracking, model production management, a model registry, and full data lineage to support your ML workflow from training straight through to production.

In 4 listsDetails

Opik

Evaluate, test, and ship LLM applications across your dev and production lifecycles.

In 16 listsDetails

Synthical

AI-powered collaborative environment for research. Find relevant papers, create collections to manage bibliography, and summarize content — all in one place

In 4 listsDetails

teeplot

Workflow tool to automatically organize data visualization output

Streamlit

App framework for Machine Learning and Data Science projects

In 10 listsDetails

Gradio

Create customizable UI components around machine learning models

In 11 listsDetails

Weights & Biases

Experiment tracking, dataset versioning, and model management

In 5 listsDetails

Optuna

Automatic hyperparameter optimization software framework

In 8 listsDetails

Ray Tune

Scalable hyperparameter tuning library

In 13 listsDetails

Apache Airflow

Platform to programmatically author, schedule, and monitor workflows

In 13 listsDetails

Prefect

Workflow management system for modern data stacks

In 11 listsDetails

Kedro

Open-source Python framework for creating reproducible, maintainable data science code

In 7 listsDetails

Hamilton

Lightweight library to author and manage reliable data transformations

In 10 listsDetails

SHAP

Game theoretic approach to explain the output of any machine learning model

In 4 listsDetails

InterpretML

InterpretML implements the Explainable Boosting Machine (EBM), a modern, fully interpretable machine learning model based on Generalized Additive Models (GAMs). This open-source package also provides visualization tools for EBMs, other glass-box models, and black-box explanations

In 8 listsDetails

LIME

Explaining the predictions of any machine learning classifier

In 5 listsDetails

flyte

Workflow automation platform for machine learning

In 5 listsDetails

dbt

Data build tool

In 7 listsDetails

zasper

Supercharged IDE for Data Science

In 3 lists

skrub

A Python library to ease preprocessing and feature engineering for tabular machine learning

In 3 lists

Glyph

Framework-agnostic TypeScript library for generating, searching, and comparing MinHash fingerprints for fast text similarity, deduplication, and retrieval.

In 2 lists

Codeflash

Ship Blazing-Fast Python Code — Every Time

In 5 listsDetails

Hugging Face

Popular open platform for sharing ML models, datasets, and collaborating on NLP and generative AI projects.

In 7 listsDetails

Chinese-Elite

An open-source project that automatically maps relationship networks by parsing public data using LLMs and visualizes it as an interactive graph.

Desbordante

An open-source data profiler specifically focused on discovery and validation of complex patterns, such as numerical association rules, differential dependencies, denial constraints, and more.

In 4 listsDetails

dna-claude-analysis

Personal genome analysis toolkit with Python scripts analyzing raw DNA data across 17 categories (health risks, ancestry, pharmacogenomics, nutrition, psychology, and more) and generating a terminal-style single-page HTML visualization.

In 7 listsDetails

RunMat

Fast MATLAB-syntax runtime with automatic CPU/GPU execution and fused array kernels.

In 5 listsDetails

Turbostream

A terminal UI for experimenting with custom rule engines and selective LLM analysis on real-time data streams, without worrying about streaming infra or backpressure.

WFGY ProblemMap

Open source “failure atlas” of 16 recurring issues in LLM and RAG pipelines, with observable symptoms and suggested fixes for data science teams.

In 4 listsDetails

Deploybase

Track real-time GPU and LLM pricing across all cloud and inference providers.

DeepAnalyze

An agentic LLM for autonomous data science, which can autonomously complete a wide range of data science tasks without human intervention.

In 8 listsDetails

Disco

Superhuman exploratory data analysis. Finds the feature interactions and subgroup effects in tabular data that LLMs and manual exploration miss — with p-values, effect sizes, and literature citations. Free for public data.

AI for Database

Chat with your database in natural language — no SQL needed. Get instant insights, build self-refreshing dashboards, and trigger automated workflows based on database changes.

In 4 listsDetails

DeepAlpha

AI crypto trading framework using LightGBM + XGBoost ensemble with 72 ML features. 70.9% walk-forward validated accuracy on out-of-sample data. Supports Bybit and Binance. MIT licensed, available on PyPI.

In 3 lists

Future AGI

Open-source platform to simulate, evaluate, trace, guardrail, route, and optimize LLM and AI agent apps in one feedback loop, so agents don't just get monitored, they self-improve. Self-hostable. Apache-2.0.

In 6 listsDetails

ipynbtopdf

Browser-based Jupyter notebook viewer and exporter that converts .ipynb notebooks to PDF, HTML, and Python without installing Python or TeX.

Literature and Media >Books

Data Science From Scratch: First Principles with Python

Artificial Intelligence with Python - Tutorialspoint

Machine Learning from Scratch

Probabilistic Machine Learning: An Introduction

by Kevin Patrick Murphy

In 2 lists

How to Lead in Data Science

Early Access

Fighting Churn With Data

In 2 lists

Data Science at Scale with Python and Dask

Python Data Science Handbook

by Jake VanderPlas

In 2 lists

The Data Science Handbook: Advice and Insights from 25 Amazing Data Scientists

Think Like a Data Scientist

Introducing Data Science

Practical Data Science with R

Everyday Data Science

& (cheaper PDF version)

Exploring Data Science

free eBook sampler

Exploring the Data Jungle

free eBook sampler

Classic Computer Science Problems in Python

Math for Programmers

Early access

R in Action, Third Edition

Early Access

Data Science Bookcamp

Early access

In 2 lists

Data Science Thinking: The Next Scientific, Technological and Economic Revolution

Applied Data Science: Lessons Learned for the Data-Driven Business

The Data Science Handbook

Essential Natural Language Processing

Early access

Mining Massive Datasets

free e-book comprehended by an online course

Pandas in Action

Early access

Genetic Algorithms and Genetic Programming

Advances in Evolutionary Algorithms

Free Download

Genetic Programming: New Approaches and Successful Applications

Free Download

Evolutionary Algorithms

Free Download

Advances in Genetic Programming, Vol. 3

Free Download

Genetic Algorithms and Evolutionary Computation

Free Download

Convex Optimization

Convex Optimization book by Stephen Boyd - Free Download

In 2 lists

Data Analysis with Python and PySpark

Early Access

In 2 lists

R for Data Science

This book will teach you how to do data science with R. You will learn how to get your data into R, get it into the most useful structure, transform it, visualize it and model it. Exercise Solutions Authors: Garrett Grolemund and Hadley Wickham.

In 3 lists

Build a Career in Data Science

Machine Learning Bookcamp

Early access

Hands-On Machine Learning with Scikit-Learn, Keras, and TensorFlow, 2nd Edition

Effective Data Science Infrastructure

In 2 lists

Practical MLOps: How to Get Ready for Production Models

Regression, a Friendly guide

Early Access

Streaming Systems: The What, Where, When, and How of Large-Scale Data Processing

Data Science at the Command Line: Facing the Future with Time-Tested Tools

Machine Learning with Python - Tutorialspoint

Deep Learning

Mathematical foundations by Ian Goodfellow, Yoshua Bengio, and Aaron Courville.

In 3 lists

Designing Cloud Data Platforms

Early Access

An Introduction to Statistical Learning with Applications in R

Gareth James, Daniela Witten, Trevor Hastie, Robert Tibshirani.

In 3 lists

The Elements of Statistical Learning: Data Mining, Inference, and Prediction

by Trevor Hastie, Robert Tibshirani, and Jerome Friedman

In 2 lists

Deep Learning with PyTorch

Neural Networks and Deep Learning

Deep Learning Cookbook

Introduction to Machine Learning with Python

Artificial Intelligence: Foundations of Computational Agents, 2nd Edition

Free HTML version

The Quest for Artificial Intelligence: A History of Ideas and Achievements

Free Download

Graph Algorithms for Data Science

Early Access

In 2 lists

Data Mesh in Action

Early Access

Julia for Data Analysis

Early Access

Regular Expression Puzzles and AI Coding Assistants

by David Mertz

Dive into Deep Learning

Interactive deep learning book with code implementations

In 4 listsDetails

Data for All

Interpretable Machine Learning: A Guide for Making Black Box Models Explainable

Free GitHub version

In 2 lists

Foundations of Data Science

Free Download

Comet for DataScience: Enhance your ability to manage and optimize the life cycle of your data science project

Software Engineering for Data Scientists

Early Access

In 2 lists

Julia for Data Science

Early Access

Machine Learning For Absolute Beginners

Unifying Business, Data, and Code: Designing Data Products with JSON Schema

Grokking Bayes

Machine Learning Q and AI

JavaScript for Data Science

Free html page

Angewandte Data Science

German book about applied data science

The Math Behind Artificial Intelligence

A free FreeCodeCamp book teaching the math behind AI in plain English from an engineering point of view.

In 4 listsDetails

Executive Data Science

A high-level guide to managing data science teams and projects.

Introduction to Modern Statistics

A modern, open-access textbook on statistics with a heavy focus on data science applications.

The Art of Data Science

Focuses on the "art" of data analysis, how to ask the right questions and refine them.

eBook sale - Save up to 45% on eBooks!

In 2 lists

Causal Machine Learning

Managing ML Projects

Causal Inference for Data Science

Literature and Media >Journals, Publications and Magazines

ICML

International Conference on Machine Learning

GECCO

The Genetic and Evolutionary Computation Conference (GECCO)

epjdatascience

In 2 lists

Journal of Data Science

an international journal devoted to applications of statistical methods at large

Big Data Research

Journal of Big Data

Big Data & Society

In 2 lists

Data Science Journal

datatau.com/news

Like Hacker News, but for data

Data Science Trello Board

Medium Data Science Topic

Data Science related publications on medium

Towards Data Science Genetic Algorithm Topic

Genetic Algorithm related Publications towards Data Science

Maxim AI

. Tool for AI Agent Simulation, Evaluation & Observability.

In 8 listsDetails

8bitconcepts

AI industry research and analysis with papers on AI pricing, enterprise adoption, and evaluation frameworks.

Literature and Media >Newsletters

AI Weekly

Curated AI intelligence briefing from industry leaders covering models, funding, policy, and applications. 3x/week since 2017, 40K+ subscribers.

DataTalks.Club

. A weekly newsletter about data-related things. Archive.

In 2 lists

The Analytics Engineering Roundup

. A newsletter about data science. Archive.

Techpresso

. A free daily newsletter covering the most impactful developments in AI, ML, and tech. Archive.

DiamantAI

. Practical AI engineering and generative AI explained simply: RAG, agents, and LLM application patterns for builders.

Bamboo Weekly

Weekly pandas exercises based on current events and real-world public data, with fully worked solutions. Issues older than two years are free, as are the first two questions + answers in current issues. Archive.

Literature and Media >Mailing lists

Working Group - Research Software Engineering in the Digital Humanities

. This is the mailing list for the Research Software Engineering in the Digital Humanities (DH-RSE) working group.

Literature and Media >Bloggers

Wes McKinney

Wes McKinney Archives.

Matthew Russell

Mining The Social Web.

Greg Reda

Greg Reda Personal Blog

Julia Evans

Recurse Center alumna

In 3 lists

Hakan Kardas

Personal Web Page

Sean J. Taylor

Personal Web Page

Drew Conway

Personal Web Page

Hilary Mason

Personal Web Page

Noah Iliinsky

Personal Blog

Matt Harrison

Personal Blog

Vamshi Ambati

AllThings Data Sciene

Prash Chan

Tech Blog on Master Data Management And Every Buzz Surrounding It

Clare Corthell

The Open Source Data Science Masters

Datawrangling

by Peter Skomoroch. MACHINE LEARNING, DATA MINING, AND MORE

Quora Data Science

Data Science Questions and Answers from experts

Siah

a PhD student at Berkeley

Louis Dorard

a technology guy with a penchant for the web and for data, big and small

Machine Learning Mastery

about helping professional programmers confidently apply machine learning algorithms to address complex problems.

In 2 lists

Daniel Forsyth

Personal Blog

Data Science Weekly

Weekly News Blog

Revolution Analytics

Data Science Blog

R Bloggers

R Bloggers

In 3 lists

The Practical Quant

Big data

Yet Another Data Blog

Yet Another Data Blog

KD Nuggets

Data Mining, Analytics, Big Data, Data, Science not a blog a portal

In 2 lists

Meta Brown

Personal Blog

Data Scientist

is building the data scientist culture.

WhatSTheBigData

is some of, all of, or much more than the above and this blog explores its impact on information technology, the business world, government agencies, and our lives.

Tevfik Kosar

Magnus Notitia

New Data Scientist

How a Social Scientist Jumps into the World of Big Data

Harvard Data Science

Thoughts on Statistical Computing and Visualization

Data Science 101

Learning To Be A Data Scientist

Kaggle Past Solutions

DataScientistJourney

NYC Taxi Visualization Blog

Data-Mania

Data-Magnum

datascopeanalytics

Digital transformation

Data Mania Blog

The File Drawer - Chris Said's science blog

Emilio Ferrara's web page

, University of Southern California, Los Angeles, USA

In 2 lists

DataNews

Reddit TextMining

Periscopic

Hilary Parker

Data Stories

In 2 lists

Data Science Lab

Meaning of

Adventures in Data Land

Dataclysm

FlowingData

Visualization and Statistics

In 3 lists

Calculated Risk

O'reilly Learning Blog

Dominodatalab

i am trask

A Machine Learning Craftsmanship Blog

Vademecum of Practical Data Science

Handbook and recipes for data-driven solutions of real-world problems

Dataconomy

A blog on the newly emerging data economy

Springboard

A blog with resources for data science learners

Analytics Vidhya

A full-fledged website about data science and analytics study material.

In 2 lists

Occam's Razor

Focused on Web Analytics.

Data School

Data science tutorials for beginners!

Colah's Blog

Blog for understanding Neural Networks!

In 2 lists

Sebastian's Blog

Blog for NLP and transfer learning!

Distill

Dedicated to clear explanations of machine learning!

In 4 listsDetails

Chris Albon's Website

Data Science and AI notes

Andrew Carr

Data Science with Esoteric programming languages

floydhub

Blog for Evolutionary Algorithms

Jingles

Review and extract key concepts from academic papers

nbshare

Data Science notebooks

Loic Tetrel

Data science blog

Chip Huyen's Blog

ML Engineering, MLOps, and the use of ML in startups

In 2 lists

Maria Khalusova

Data science blog

Aditi Rastogi

ML,DL,Data Science blog

Santiago Basulto

Data Science with Python

Akhil Soni

ML, DL and Data Science

Akhil Soni

ML, DL and Data Science

Applied AI Blogs

In-depth articles on AI, machine learning, and data science concepts with practical applications.

In 2 lists

Scaler Blogs

Educational content on software development, AI, and career growth in tech.

Mlu github

Mlu is developed amazon to help people in ml space you can learn everything from basics here with live diagrams

Jan Oliver Rüdiger

ML, DL and Data Science - with a focus on text-/data-mining

Literature and Media >Presentations

How to Become a Data Scientist

Introduction to Data Science

Intro to Data Science for Enterprise Big Data

How to Interview a Data Scientist

How to Share Data with a Statistician

Guide to data sharing.

In 2 lists

The Science of a Great Career in Data Science

What Does a Data Scientist Do?

Building Data Start-Ups: Fast, Big, and Focused

How to win data science competitions with Deep Learning

Full-Stack Data Scientist

Literature and Media >Podcasts

AI at Home

AI Today

Adversarial Learning

Chai time Data Science

Chain of Thought

AI infrastructure and developer tools, with interviews from engineering leaders and technical founders.

In 4 listsDetails

Data Engineering Podcast

Description: A show that digs deep into databases, big data pipelines, data governance, data collection, ETL, building effective data teams, and all of the other challenges faced by technology professionals working on data management.; Frequency: Once a week; Host: Tobias Macey; Runtime: 30-60…

In 5 listsDetails

Data Science at Home

Data Science Mixer

Data Skeptic

Description: High level concepts in data science, and longer interview with researchers and practitioners; Host: Kyle Polich @DataSkeptic; Frequency: Once a week; Runtime: 15 - 45 mins, regularly ~25 mins

In 2 lists

Data Stories

In 2 lists

Datacast

DataFramed

Description: Interviews with industry experts on what data science is, what problems it tries to solve and what it looks like in practice.; Host: Hugo Bowne-Anderson @hugobowne; Frequency: Once a week; Runtime: 50 - 60 mins

In 3 lists

DataTalks.Club

Gradient Descent

Learning Machines 101

Let's Data (Brazil)

Linear Digressions

Not So Standard Deviations

O'Reilly Data Show Podcast

Partially Derivative

Superdatascience

Description: Podcast that interviews various people in the data science industry.; Host: Chris Benson @chrisbenson, Daniel Whitenack @dwhitena; Frequency: Once a week; Runtime: Alternates between ~5 and ~60 mins

In 2 lists

The Data Engineering Show

The Radical AI Podcast

What's The Point

The Analytics Engineering Podcast

How analytics engineers build and maintain data pipelines at scale.

In 2 lists

Literature and Media >YouTube Videos & Channels

What is machine learning?

Andrew Ng: Deep Learning, Self-Taught Learning and Unsupervised Feature Learning

In 2 lists

Data36 - Data Science for Beginners by Tomi Mester

Deep Learning: Intelligence from Big Data

Interview with Google's AI and Deep Learning 'Godfather' Geoffrey Hinton

Introduction to Deep Learning with Python

What is machine learning, and how does it work?

CampusX

Data School

Data Science Education

Neural Nets for Newbies by Melanie Warrick (May 2015)

Neural Networks video series by Hugo Larochelle

Interesting class about neural networks available online for free by Hugo Larochelle, yet I have watched a few of those videos.

In 2 lists

Google DeepMind co-founder Shane Legg - Machine Super Intelligence

Data Science Primer

Data Science with Genetic Algorithms

Data Science for Beginners

DataTalks.Club

Mildlyoverfitted - Tutorials on intermediate ML/DL topics

ML Street Talk - Unabashedly technical and non-commercial, so you will hear no annoying pitches.

Neural networks by 3Blue1Brown

Neural networks from scratch by Sentdex

Manning Publications YouTube channel

Ask Dr Chong: How to Lead in Data Science - Part 1

Ask Dr Chong: How to Lead in Data Science - Part 2

Ask Dr Chong: How to Lead in Data Science - Part 3

Ask Dr Chong: How to Lead in Data Science - Part 4

Ask Dr Chong: How to Lead in Data Science - Part 5

Ask Dr Chong: How to Lead in Data Science - Part 6

Regression Models: Applying simple Poisson regression

Deep Learning Architectures

Time Series Modelling and Analysis

Serrano.Academy

End to End Data Science Playlist

Introduction to Data Science - Linkedin

AI Talks

Searchable summaries and topic index for practical AI engineering talks and conference videos.

Socialize >Facebook Accounts

Data

Big Data Scientist

Data Science Day

Data Science Academy

Facebook Data Science Page

Data Science London

Data Science Technology and Corporation

Data Science - Closed Group

Center for Data Science

Big data hadoop NOSQL Hive Hbase

Analytics, Data Mining, Predictive Modeling, Artificial Intelligence

Big Data Analytics using R

Big Data Analytics with R and Hadoop

Big Data Learnings

Big Data, Data Science, Data Mining & Statistics

BigData/Hadoop Expert

Data Mining / Machine Learning / AI

Data Mining/Big Data - Social Network Ana

Vademecum of Practical Data Science

Veri Bilimi Istanbul

The Data Science Blog

Socialize >Twitter Accounts

Big Data Combine

Rapid-fire, live tryouts for data scientists seeking to monetize their models as trading strategies

Big Data Science

Big Data, Data Science, Predictive Modeling, Business Analytics, Hadoop, Decision and Operations Research.

Chris Said

Data scientist at Twitter

Clare Corthell

Dev, Design, Data Science @mattermark #hackerei

DADI Charles-Abner

#datascientist @Ekimetrics. , #machinelearning #dataviz #DynamicCharts #Hadoop #R #Python #NLP #Bitcoin #dataenthousiast

Data Science Central

Data Science Central is the industry's single resource for Big Data practitioners.

Data Science London

Data Science. Big Data. Data Hacks. Data Junkies. Data Startups. Open Data

Data Science Renee

Documenting my path from SQL Data Analyst pursuing an Engineering Master's Degree to Data Scientist

Data Science Report

Mission is to help guide & advance careers in Data Science & Analytics

Data Science Tips

Tips and Tricks for Data Scientists around the world! #datascience #bigdata

Data Vizzard

DataViz, Security, Military

DataScienceX

DJ Patil

White House Data Chief, VP @ RelateIQ.

Domino Data Lab

Drew Conway

Data nerd, hacker, student of conflict.

Erin Bartolo

Running with #BigData--enjoying a love/hate relationship with its hype. @iSchoolSU #DataScience Program Mgr.

Greg Reda

Working @ GrubHub about data and pandas

Gregory Piatetsky

KDnuggets President, Analytics/Big Data/Data Mining/Data Science expert, KDD & SIGKDD co-founder, was Chief Scientist at 2 startups, part-time philosopher.

Hadley Wickham

Chief Scientist at RStudio, and an Adjunct Professor of Statistics at the University of Auckland, Stanford University, and Rice University.

Hakan Kardas

Data Scientist

Hilary Mason

Data Scientist in Residence at @accel.

Jeff Hammerbacher

ReTweeting about data science

John Myles White

Scientist at Facebook and Julia developer. Author of Machine Learning for Hackers and Bandit Algorithms for Website Optimization. Tweets reflect my views only.

Juan Miguel Lavista

Principal Data Scientist @ Microsoft Data Science Team

Julia Evans

Hacker - Pandas - Data Analyze

Kenneth Cukier

The Economist's Data Editor and co-author of Big Data (https://www.big-data-book.com/).

Kevin Markham

Data science instructor, and founder of Data School

Kim Rees

Interactive data visualization and tools. Data flaneur.

Kirk Borne

DataScientist, PhD Astrophysicist, Top #BigData Influencer.

Luis Rei

PhD Student. Programming, Mobile, Web. Artificial Intelligence, Intelligent Robotics Machine Learning, Data Mining, Natural Language Processing, Data Science.

Matt Harrison

Opinions of full-stack Python guy, author, instructor, currently playing Data Scientist. Occasional fathering, husbanding, organic gardening.

Matthew Russell

Mining the Social Web.

Mert Nuhoğlu

Data Scientist at BizQualify, Developer

Monica Rogati

Data @ Jawbone. Turned data into stories & products at LinkedIn. Text mining, applied machine learning, recommender systems. Ex-gamer, ex-machine coder; namer.

Noah Iliinsky

Visualization & interaction designer. Practical cyclist. Author of vis books: https://www.oreilly.com/pub/au/4419

Paul Miller

Cloud Computing/ Big Data/ Open Data Analyst & Consultant. Writer, Speaker & Moderator. Gigaom Research Analyst.

Peter Skomoroch

Creating intelligent systems to automate tasks & improve decisions. Entrepreneur, ex-Principal Data Scientist @LinkedIn. Machine Learning, ProductRei, Networks

Prash Chan

Solution Architect @ IBM, Master Data Management, Data Quality & Data Governance Blogger. Data Science, Hadoop, Big Data & Cloud.

Quora Data Science

Quora's data science topic

R-Bloggers

Tweet blog posts from the R blogosphere, data science conferences, and (!) open jobs for data scientists.

Rand Hindi

Randy Olson

Computer scientist researching artificial intelligence. Data tinkerer. Community leader for @DataIsBeautiful. #OpenScience advocate.

Recep Erol

Data Science geek @ UALR

Ryan Orban

Data scientist, genetic origamist, hardware aficionado

Sean J. Taylor

Social Scientist. Hacker. Facebook Data Science Team. Keywords: Experiments, Causal Inference, Statistics, Machine Learning, Economics.

Silvia K. Spiva

#DataScience at Cisco

Harsh B. Gupta

Data Scientist at BBVA Compass

Spencer Nelson

Data nerd

Talha Oz

Enjoys ABM, SNA, DM, ML, NLP, HI, Python, Java. Top percentile Kaggler/data scientist

Tasos Skarlatidis

Complex Event Processing, Big Data, Artificial Intelligence and Machine Learning. Passionate about programming and open-source.

Terry Timko

InfoGov; Bigdata; Data as a Service; Data Science; Open, Social & Business Data Convergence

Tony Baer

IT analyst with Ovum covering Big Data & data management with some systems engineering thrown in.

Tony Ojeda

Data Scientist , Author , Entrepreneur. Co-founder @DataCommunityDC. Founder @DistrictDataLab. #DataScience #BigData #DataDC

Vamshi Ambati

Data Science @ PayPal. #NLP, #machinelearning; PhD, Carnegie Mellon alumni (Blog: https://allthingsds.wordpress.com )

Wes McKinney

Pandas (Python Data Analysis library).

WileyEd

Senior Manager - @Seagate Big Data Analytics @McKinsey Alum #BigData + #Analytics Evangelist #Hadoop, #Cloud, #Digital, & #R Enthusiast

WNYC Data News Team

The data news crew at @WNYC. Practicing data-driven journalism, making it visual, and showing our work.

Alexey Grigorev

Data science author

İlker Arslan

Data science author. Shares mostly about Julia programming

INEVITABLE

AI & Data Science Start-up Company based in England, UK

Jan Oliver Rüdiger

ML, DL and Data Science - with a focus on text-/data-mining

Socialize >Telegram Channels

Open Data Science

First Telegram Data Science channel. Covering all technical and popular staff about anything related to Data Science: AI, Big Data, Machine Learning, Statistics, general Math and the applications of former.

Loss function porn

Beautiful posts on DS/ML theme with video or graphic visualization.

Machinelearning

Daily ML news.

Socialize >Slack Communities

DataTalks.Club

. A weekly newsletter about data-related things. Archive.

In 2 lists

Socialize >GitHub Groups

Berkeley Institute for Data Science

Socialize >Data Science Competitions

Kaggle

kaggle has built-in free jupyter notebook.; One can also connect to Google BigQuery to access big data.

In 5 listsDetails

DrivenData

Participate in data science competitions and help organizations.

In 2 lists

Analytics Vidhya

InnoCentive

In 2 lists

Realtime deployment

Tutorial on Python time-series model deployment.

Fun >Datasets

Academic Torrents

ADS-B Exchange

Specific datasets for aircraft and Automatic Dependent Surveillance-Broadcast (ADS-B) sources.

Chinese Tea Dataset

Curated open dataset of 100+ Chinese teas with category, origin, caffeine level, flavor notes, oxidation, and brewing parameters. Available as JSON and CSV.

College ROI Dataset

Lifetime return-on-investment estimates for ~30K US bachelor's programs across 1,775 institutions, built from FREOPP, IPEDS, and BEA regional price data. 5 CSVs with data dictionary, CC BY 4.0, Zenodo DOI.

AI Displacement Tracker

Structured dataset tracking 92 AI-attributed workforce reduction events affecting 453,748 workers across 12 countries and 11 sectors. JSON and CSV formats. CC-BY-4.0 licensed.

Packrift Packaging Optimization Benchmark Corpus

Public packaging product dataset generated from 1,000 exact-spec SKU records, with downloadable CSV and JSON files for ecommerce fulfillment and warehouse analysis.

Pokemon Card Centering Measurements

320 measured PSA-style centering annotations (left/right and top/bottom border percentages, tilt) across 302 real eBay-listed Pokemon cards. CSV, CC BY 4.0, Zenodo DOI.

Pokemon Card Sold-Price Reference by Grade

Median sold price by grade (raw, PSA 9, PSA 10) for 486 Pokemon cards, with sample size and confidence flag per card. CSV, CC BY 4.0, Zenodo DOI.

Evidaxis Momentum Snapshots

Weekly snapshots of public development and citation activity for open-source and research-native AI systems, content-addressed and byte-reproducible from public inputs. JSON and CSV per snapshot date, CC0, DOI 10.5281/zenodo.21076011.

hadoopilluminated.com

data.gov

The home of the U.S. Government's open data

United States Census Bureau

enigma.com

Navigate the world of public data - Quickly search and analyze billions of public records published by governments, companies and organizations.

In 2 lists

datahub.io

aws.amazon.com/datasets

Amazon public datasets.

In 3 lists

datacite.org

The official portal for European data

NASDAQ:DATA

Nasdaq Data Link A premier source for financial, economic and alternative datasets.

In 2 lists

Congressional Stock Brain

Free AI-powered tool that scores U.S. congressional STOCK Act trade disclosures by significance. Machine-scored signals from 537 lawmakers's public trade filings.

In 3 lists

figshare.com

(Storage, Lookup): Data sharing and storage

In 2 lists

GeoLite Legacy Downloadable Databases

Hugging Face Datasets

the central index for modern NLP datasets, with versioned, streamable loaders.

In 5 listsDetails

Japan Neighborhoods

English dataset of Tokyo crime statistics across 5,078 neighborhoods × 7 years (36,222 records, 2018-2024), sourced from Tokyo Metropolitan Police open data. Includes interactive crime map, safety grading, and cost-of-living index. CC BY licensed.

In 2 lists

The Quiet-Broke Index

A 30-metro composite ranking of how much of a $400K household income gets consumed by housing, taxes, childcare, healthcare, and transport. Open methodology, free, no email gate.

In 2 lists

Crime Brasil

Open-data platform for Brazilian crime statistics. Neighborhood-level in Rio Grande do Sul (2.99M incidents across 79,024 neighborhoods, 2022–2025), municipality-level for MG and RJ, plus national PRF highway and DATASUS interpersonal-violence data. Free REST API, CSV/Parquet, daily updates, CC BY…

In 5 listsDetails

US Truck-Involved Fatal Crashes (FARS) 2018-2024

Filtered subset of NHTSA Fatality Analysis Reporting System covering 33,898 fatal crashes involving medium and heavy commercial trucks across all 50 US states, 2018-2024. Includes interactive Vision Zero Report Card comparing 19 cities, reproducible Python pipeline on GitHub, and HuggingFace…

State of Peptides 2026

Structured reference dataset of 156 peptide and peptide-adjacent compounds, each with a regulatory status bucket, category, route, half-life, molecular weight, CAS number, reference count, and PubChem/DrugBank/Wikidata IDs. CSV and JSON, no login, CC BY 4.0.

Quora's Big Datasets Answer

Kaggle Datasets

Extensive collection of datasets for practice in data analysis.

In 3 lists

A Deep Catalog of Human Genetic Variation

A community-curated database of well-known people, places, and things

Large scale knowledge base originally stated by Metaweb. Later aquired by Google and used in Google Knowledge Graph.

In 3 lists

Google Public Data

In 2 lists

World Bank Data

Free and open access to global development data by The World Bank.

In 3 lists

NYC Taxi Visualization Blog

Open Data Philly

Connecting people with data for Philadelphia

grouplens.org

Sample movie (with ratings), book and wiki datasets

UC Irvine Machine Learning Repository

contains data sets good for machine learning

research-quality data sets

by Hilary Mason

National Centers for Environmental Information

(Lookup): Weather, climate, coasts, oceans, and geophysics etc

In 2 lists

ClimateData.us

(related: U.S. Climate Resilience Toolkit)

r/datasets

Datasets for Data Mining, Analytics and Knowledge Discovery.

In 3 lists

MapLight

provides a variety of data free of charge for uses that are freely available to the general public. Click on a data set below to learn more

GHDx

Institute for Health Metrics and Evaluation - a catalog of health and demographic datasets from around the world and including IHME results

St. Louis Federal Reserve Economic Data - FRED

New Zealand Institute of Economic Research – Data1850

Open Data Sources

Collection of various open data sources.

In 2 lists

UNICEF Data

undata

Data from the UN

In 3 lists

NASA SocioEconomic Data and Applications Center - SEDAC

The GDELT Project

The GDELT Project monitors the world's broadcast, print, and web news from nearly every corner of every country in over 100 languages and identifies the people, locations, organizations, themes, sources, emotions, counts, quotes, images and events driving our global society every second of every…

In 3 lists

Sweden, Statistics

StackExchange Data Explorer

an open source tool for running arbitrary queries against public data from the Stack Exchange network.

San Fransisco Government Open Data

IBM Asset Dataset

Open data Index

Public Git Archive

6 TB of Git repositories from GitHub.

In 2 lists

GHTorrent

Microsoft Research Open Data

Open Government Data Platform India

Google Dataset Search (beta)

Dataset Search is a search engine for datasets. Using a simple keyword search, users can discover datasets hosted in thousands of repositories across the Web.

In 6 listsDetails

NAYN.CO Turkish News with categories

Covid-19

Covid-19 Google

Enron Email Dataset

5000 Images of Clothes

IBB Open Portal

The Humanitarian Data Exchange

250k+ Job Postings

An expanding dataset of historical job postings from Luxembourg from 2020 to today. Free with 250k+ job postings hosted on AWS Data Exchange.

FinancialData.Net

Financial datasets (stock market data, financial statements, sustainability data, and more).

In 2 lists

HDD Price Index

Daily open dataset of the cheapest new internal 3.5" SATA hard-drive price per terabyte (USD/TB) by capacity tier on Amazon US, with a historical time series. CSV, JSON and JSONL, no login, CC BY 4.0.

BDE Score

AI-powered multi-market stock analysis with transparent BDE scoring across 73 stocks (US/HK/A-share). EU AI Act Art.50 compliant. MIT license.

In 4 listsDetails

notesjor corpus-collection

Free corpora (over 6 billion tokens) mostly German (both historically and in contemporary German).

CLARIN-Repository

CLARIN is a European repository for scientific datasets.

GBIF

Global Biodiversity Information Facility: 2.4B+ species occurrence records. Free, open API for ecological modeling and ML research.

In 3 lists

FAOSTAT

UN FAO statistics on food production, trade, land use, and emissions for 245+ countries. Free API and bulk download.

Movebank

Free platform archiving 6B+ animal movement records from GPS and satellite telemetry. Open REST API, useful for spatiotemporal modeling and trajectory ML.

Encyclopedia of Life

Open structured data on 1.9M+ species, including traits, classification, and media. Free API and bulk downloads for biodiversity and species-classification tasks.

FirstData

The world's most comprehensive authoritative data source knowledge base. 210+ curated sources from governments, international organizations, and research institutions. MCP integration for AI agents. MIT licensed.

In 2 lists

latamdata-py

Python package for one-line access to 38 open research datasets from Latin America (health, neuroscience, mental health, economics). pip install latamdata-py.

ZipCheckup

Free ZIP-level environmental safety data for 42,000+ US ZIP codes: water quality, air quality, PFAS contamination, radon, lead, flood risk, and 11 more verticals. Public REST API, npm/PyPI packages, CC BY 4.0.

In 3 lists

Helium

Real-time news corpus with structured bias features across 15+ dimensions (3.2M+ articles, 5,000+ sources), live financial market data (stocks, ETFs, crypto) with AI-generated analysis, ML options pricing with probability metrics and full Greeks, historical options chain data for quantitative…

In 3 lists

Verified Supplement Evidence

Evidence-graded dietary-supplement dataset covering dosing, bioavailability by form, drug-nutrient interactions, NHANES deficiency prevalence, FDA FAERS adverse-event signals, and cost-per-effective-dose, with every clinical claim citing a PubMed PMID. CC BY 4.0, DOI 10.57967/hf/9356.

WhatFontIs-Bench

Synthetic benchmark for font family identification with 11,995 images of words set in 600 known fonts, annotated with word and per-letter boxes.

Fun >Comics

Comic compilation

Cartoons

Data Science Cartoons

Data Science: The XKCD Edition

See category
94

Table of Contents

hesreallyhim/awesome-claude-code

A hand-picked collection of the finest of resources for the most awesome of agents, Claude Code, the undisputed champion of coding companions, from the unstoppable team…

Fresh★ 55k202 entriesPushed today
94

Awesome Agent Skills

VoltAgent/awesome-agent-skills

A curated collection of 1000+ agent skills from official dev teams and the community, compatible with Claude Code, Codex, Gemini CLI, Cursor, and more.

Fresh★ 35k839 entriesPushed today
93

Awesome Machine Learning

josephmisiti/awesome-machine-learning

A curated list of awesome Machine Learning frameworks, libraries and software.

Fresh★ 74k1188 entriesPushed 7 days ago
92

Awesome Production Machine Learning

EthicalML/awesome-production-machine-learning

A curated list of awesome open source libraries to deploy, monitor, version and scale your machine learning

Fresh★ 21k519 entriesPushed 3 days ago
91

Static Analysis

analysis-tools-dev/static-analysis

⚙️ A curated list of static analysis (SAST) tools and linters for all programming languages, config files, build tools, and more. The focus is on tools which improve…

Fresh★ 15k528 entriesPushed 8 days ago
90

Awesome LangChain

kyrolabs/awesome-langchain

😎 Awesome list of tools and projects with the awesome LangChain framework

Fresh★ 9.6k216 entriesPushed 6 days ago