Data Science For Beginners
Microsoft are pleased to offer a 10-week, 20-lesson curriculum all about Data Science.
:memo: An awesome Data Science repository to learn and apply for real world problems.
This page lists names, links and short descriptions. The original list on GitHub is the source and belongs to its authors.
Microsoft are pleased to offer a 10-week, 20-lesson curriculum all about Data Science.
Data scientists combine entrepreneurship with patience, the willingness to build data products incrementally, the ability to explore, and the ability to iterate over a solution. They are inherently interdisciplinary. They can tackle all aspects of a problem, from initial data collection and data…
Data Science is a combination of a number of aspects of Data such as Technology, Algorithm development, and data interference to study the data, analyse it, and find innovative solutions to difficult problems. Basically Data Science is all about Analysing data and driving for business growth by…
Data scientists today are akin to Wall Street “quants” of the 1980s and 1990s. In those days people with backgrounds in physics and math streamed to investment banks and hedge funds, where they could devise entirely new algorithms and data strategies. Then a variety of universities developed…
Data science is an interdisciplinary field that uses scientific methods, processes, algorithms and systems to extract knowledge and insights from many structural and unstructured data. Data science is related to data mining, machine learning and big data.
Data scientists are big data wranglers, gathering and analyzing large sets of structured and unstructured data. A data scientist’s role combines computer science, statistics, and mathematics. They analyze, process, and model data then interpret the results to create actionable plans for companies…
The story of how data scientists became sexy is mostly the story of the coupling of the mature discipline of statistics with a very young one--computer science. The term “Data Science” has emerged only recently to specifically designate a new profession that is expected to make sense of the vast…
Data scientists concentrate on making sense of data through exploratory analysis, statistics, and models. Software developers apply a separate set of knowledge with different tools. Although their focus may seem unrelated, data science teams can benefit from adopting software development best…
Data science is an excellent career choice in today’s data-driven world where approx 328.77 million terabytes of data are generated daily. And this number is only increasing day by day, which in turn increases the demand for skilled data scientists who can utilize this data to drive business growth.
_Data science is one of the most in-demand careers today. With businesses increasingly relying on data to make decisions, the need for skilled data scientists has grown rapidly. Whether it’s tech companies, healthcare organizations, or even government institutions, data scientists play a crucial…
is by far the most popular language in science, due in no small part to the ease at which it can be used and the vibrant ecosystem of user-generated packages. To install packages, there are two main methods: Pip (invoked as pip install), the package manager that comes bundled with Python, and…
Production-ready AI agent development kit for Rust with model-agnostic design (Gemini, OpenAI, Anthropic), multiple agent types (LLM, Graph, Workflow), MCP support, and built-in telemetry.
Agent framework for chatting with data, turning natural language into SQL, transformation pipelines and visualizations. Outputs are declarative specs that can be inspected, edited, reopened in a notebook or composed into a dashboard.
MCP server providing 13 data tools for AI agents: real-time crypto prices, IP geolocation, DNS lookups, web scraping to markdown, code execution, and screenshots. One API key for 40+ services.
61 production-ready AI API tools for data science workflows: code analysis, web scraping, NLP, image generation, crypto data, and search. REST API and MCP protocol support. GitHub
Search engine for AI agents that indexes 9,000+ AI tools and APIs, scoring each on agentic readiness (llms.txt, OpenAPI, MCP, ai-plugin.json). REST API and MCP server for programmatic tool discovery. GitHub
AI crypto trading framework using LightGBM + XGBoost ensemble with 72 ML features. 70.9% walk-forward validated accuracy on out-of-sample data. Supports Bybit and Binance. MIT licensed, available on PyPI.
Local AI agent for generating publication-ready scientific papers with real arXiv citations, IMRaD structure, and tribunal scoring. Runs 100% offline via Ollama with 4B-9B models. MIT licensed. HuggingFace
Open-source LLM and agent evaluation framework with 50+ metrics, LLM-as-Judge augmentation, and guardrail scanners (jailbreak, PII, prompt-injection). Useful for scoring RAG outputs, agent trajectories, and function-calling behavior in data-science workflows.
Open-source platform that records real AI agent runs, replays them against changes, and evaluates outcomes before deployment.
Read-only social research agent that lets Jev choose bounded Instagram, TikTok, and LinkedIn operations, runs them through the local socai CLI in Chrome, and preserves source-linked evidence beside a cited report.
Open-source trusted-host experiment runner for historical agent tasks, supplied coding prompts, and workflows. Runs independent attempts with retained outputs, and compares models, harnesses, and configurations with different checks or judges later. MIT licensed.
MCP server that gives AI agents access to a database of scientific papers built from raw experimental data extracted from full-text studies. Returns 25+ structured fields per paper including methods, results, sample sizes, and quality scores. GitHub
Open-source Python library and MCP server to benchmark document chunking strategies for RAG, score retrieval quality, and recommend configurations for a corpus.
Daily-updated skill and CLI for deterministic retrieval across arXiv, PubMed/PMC, and supported US policy corpora.
x402 payment gateway with 23 Research & Reference endpoints for AI agents: Wikipedia, arXiv, PubMed, Wikidata, academic citation lookup, entity extraction, and more. Pay-per-call in USDC on Base & Solana — no API keys or subscriptions. Also serves 150+ endpoints across 39 categories including…
AI literature search, document translation, and deep-research workspace for researchers.
Sim Studio's interface is a lightweight, intuitive way to quickly build and deploy LLMs that connect with your favorite tools.
A Medical Benchmark & EHR Simulation Platform
you can run on the browser with IPython.
A weekly data project aimed at the R ecosystem.
Cheatsheets for data science.
Common PySpark patterns.
LiveVideo tutorial that covers machine learning, Tensorflow, artificial intelligence, and neural networks.
Free cross-platform search engine indexing 50,000+ tutorials from Udemy, Skillshare, Pluralsight, and other major learning platforms across 45+ categories.
Tutorial on Python time-series model deployment.
A collection of resume examples and tips tailored for data scientists.
Interview practice with SQL query execution, Python, and data modeling exercises.
Interactive calculator that visualizes the step-by-step manual math behind machine learning algorithms for exam prep.
A developer handbook on designing and building effective AI agents.
A straightforward method for training your LLM, from downloading data to generating text.
Open Source Society University
Roadmap to becoming an Artificial Intelligence Expert
Convex Optimization (basics of convex analysis; least-squares, linear and quadratic programs, semidefinite programming, minimax, extremal volume, and other problems; optimality conditions, duality theory...)
Introduction to machine learning covering basic theory, algorithms and applications
Learn about Data Science, Machine Learning, Python etc
Learn how to monitor and root-cause production ML issues.
Free Course and Certification for building an end-to-end machine using W&B
This course is designed to empower beginners with the essential skills to excel in today's data-driven world. The comprehensive curriculum will give you a solid foundation in statistics, programming, data visualization, and machine learning.
Slides, scripts and materials for the Machine Learning in Finance course at NYU Tandon, 2022.
A hands-on course to train and deploy a serverless API that predicts crypto prices.
Learn to build modern software with LLMs using the newest tools and techniques in the field.
Learn to prompt cutting-edge computer vision models with natural language, coordinate points, bounding boxes, segmentation masks, and even other images in this free course from DeepLearning.AI.
Free resources and learn what data science is and how it’s used in different industries.
A free video series by Andrej Karpathy covering neural networks from scratch — backpropagation, makemore, GPT, and more.
, Video Lectures Summer School Montreal
30 weeks
Fei-Fei Li, Andrej Karphaty, Justin Johnson (Stanford University)
New series of 5 Deep Learning courses by Andrew Ng, now with Python rather than Matlab/Octave, and which leads to a specialization certificate.
Teaches how to use machine learning to understand and manipulate human language. It requires a working knowledge of machine learning, intermediate Python experience including DL frameworks & proficiency in calculus, linear algebra, & statistics.
Interactive courses for learning data analysis.
Linear Algebra course by Gilbert Strang
is an intermediate/advanced level specialization focused on Recommender System on the Coursera platform.
Getting Started with Python for Data Science
Professional courses in data analysis, statistics, and machine learning fundamentals.
course material on text-mining / corpus-linguistics in German funded by the federal state of North Rhine-Westphalia
course material: programming in python in German for digital humanities - funded by the federal state of North Rhine-Westphalia
Short lessons with hands-on coding exercises and spaced repetition, covering Python, PyTorch, math for ML, ML foundations, NLP, and computer vision.
A collection of online data science and analytics certificate, postgraduate, and degree programs.
is an approach to model the relationship between two variables by fitting a linear equation to observed data. One variable is considered to be an explanatory variable, and the other is considered to be a dependent variable.
Easily plottable and understandable classification.
The prototypical clustering method.
The tree provides an interpretation.
, Python Code
scikit-learn clustering algorithms.
is a machine learning algorithm that is used solved calssification problems. It's based on applying Bayes' theorem with strong independence assumptions between the features.
is a fast, scalable, high performance Gradient Boosting on Decision Trees library, used for ranking, classification, regression and other machine learning tasks for Python, R, Java, C++. Supports computation on CPU and GPU.
A Python module for machine learning built on top of SciPy.
Multi-label classification for python.
Highly interpretable classifiers for scikit learn.
Feature selection repository in Python.
A scikit-learn-compatible Python implementation of ReBATE, a suite of Relief-based feature selection algorithms for Machine Learning.
Sequence classification toolkit for Python.
Python package for Bayesian Machine Learning with scikit-learn API.
A scikit-learn-inspired API for CRFsuite.
Use evolutionary algorithms instead of gridsearch in scikit-learn.
Model evaluation made easy: plots, tables, and markdown reports.
A collection of algorithms for image processing in Python.
Swarm Intelligence in Python (Genetic Algorithm, Particle Swarm Optimization, Simulated Annealing, Ant Colony Algorithm, Immune Algorithm, Artificial Fish Swarm Algorithm in Python)
Post-hoc tests for statistical analysis of data.
Memory-efficient FastText variant with exact trie n-gram IDs, structure-aware row sharing, and mmap serving for large-vocabulary NLP.
Simple structured learning framework for Python.
A high performance, easy-to-use, and scalable machine learning package, which can be used to solve large-scale machine learning problems. xLearn is especially useful for solving machine learning problems on large-scale sparse data, which is very common in Internet services such as online…
is a suite of libraries that implement machine learning algorithms and mathematical primitives functions that share compatible APIs with other RAPIDS projects. cuML enables data scientists, researchers, and software engineers to run traditional tabular ML tasks on GPUs without going into the…
Uplift modeling and causal inference with machine learning algorithms.
A library of extension and helper modules for Python's data analysis and machine learning libraries.
PySpark + scikit-learn = Sparkit-learn.
50%+ Faster, 50%+ less RAM usage, GPU support re-written Sklearn, Statsmodels.
zap: - A toolkit for making real world machine learning and data analysis applications in C++. [Boost] website
A Java port of SciPy's signal processing module, offering filters, transformations, and other scientific computing utilities.
Implementation of the rulefit.
A Python library for generalized additive models with built-in smoothing and regularization.
Validation & testing of machine learning models and data during model development, deployment, and production. This includes checks and suites related to various types of issues, such as model performance, data integrity, distribution mismatches, and more.
Scalable, Portable and Distributed Gradient Boosting (GBDT, GBRT or GBM) Library, for Python, R, Java, Scala, C++ and more. Runs on single machine, Hadoop, Spark, Flink and DataFlow. [Apache2]
Microsoft's fast, distributed, high performance gradient boosting (GBDT, GBRT, GBM or MART) framework based on decision tree algorithms, used for ranking, classification and many other machine learning tasks.
General purpose gradient boosting on decision trees library with categorical features support out of the box. It is easy to install, contains fast inference implementation and supports CPU and GPU (even multi-GPU) computation.
A gradient boosting machine that doesn't need hyperparameter optimization, with a simple budget parameter to control model complexity.
| Python | - Composable transformations of Python+NumPy programs: differentiate, vectorize, JIT to GPU/TPU, and more
Scikit-learn native toolkit for nonprofit fundraising analytics: leakage-safe donor propensity, lapse, planned-giving, wealth-screening and revenue-forecasting estimators.
(label: good first issue) PyTorch is an open source machine learning library based on the Torch library, used for applications such as computer vision and natural language processing.
GPU and multi-GPU dimensionality reduction with a scikit-learn-compatible API.
PyTorch's official computer vision library with 50+ pre-trained model architectures including ResNet, EfficientNet, Vision Transformers (ViT), ConvNeXt, and more. The de facto standard model zoo for PyTorch computer vision. BSD-3-Clause licensed.
:Torchtext是一个非常好用的库,可以帮助我们很好的解决文本的预处理问题。此github存储库包含两部分:; torchText.data:文本的通用数据加载器、抽象和迭代器(包括词汇和词向量); torchText.datasets:通用NLP数据集的预训练加载程序 我们只需要通过pip install torchtext安装好torchtext后,便可以开始体验Torchtext 的种种便捷之处。
A set of tools and building blocks for audio and speech processing, designed to accelerate the development and deployment of machine learning applications in these domains. It offers GPU-compatible, differentiable, and production-ready components, making it valuable for integrating audio…
High-level library for training and evaluating neural networks in PyTorch with an engine, events & handlers system for maximum flexibility. BSD-3-Clause licensed.
A Keras-like framework for PyTorch that handles much of the boilerplating code needed to train neural networks.
scikit-learn compatible neural network library that wraps PyTorch. Seamlessly integrate PyTorch models with scikit-learn pipelines, grid search, and cross-validation.
Python package facilitating the use of Bayesian Deep Learning methods with Variational Inference for PyTorch.
Graph neural network library for PyTorch enabling molecular modeling, materials discovery, protein interaction networks, and scientific knowledge graph learning (23.7k+ stars)
A highly efficient and modular implementation of Gaussian Processes in PyTorch.
Deep universal probabilistic programming with Python and PyTorch.
High-level utils for PyTorch DL & RL research. It was developed with a focus on reproducibility, fast experimentation and code/ideas reusing. Being able to research/develop something new, rather than write another regular train loop.
PyTorch implementation of YOLOv3, YOLOv3-SPP, and YOLOv3-tiny for real-time object detection with training, validation, inference, and multi-format export.
Ultralytics YOLOv5 in PyTorch for object detection, instance segmentation, classification, training, and export.
Ultralytics YOLO27, YOLO26, YOLO11, YOLOv8 — object detection, instance segmentation, semantic segmentation, image classification, pose estimation, object tracking
PyTorch-native library for building, training and teaching transformer language models, with architectures written as ordinary nn.Modules.
How to use the Hexagon Delegate to speed up model inference on mobile and edge devices. Also see blog post Accelerating TensorFlow Lite on Qualcomm Hexagon DSPs.
Deep learning and reinforcement learning library for researchers and engineers
Sonnet is DeepMind's library built on top of TensorFlow for building complex neural networks.
TensorFlow Reinforcement Learning.
MLOps Tools For Managing & Orchestrating The Machine Learning LifeCycle. Reproducible and scalable machine learning workflows on Kubernetes with experiment tracking, model management, and pipeline orchestration. Apache 2.0 licensed.
Deploy TensorFlow graphs for fast evaluation and export to TensorFlow-less environments running numpy.
TensorFlow ROCm port.
Deep learning with dynamic computation graphs in TensorFlow.
A high-level framework for TensorFlow.
Model Parallelism Made Easier.
Low-code framework for building custom LLMs and deep neural networks. Declarative YAML configuration for training state-of-the-art models with PEFT/LoRA, 4-bit quantization, distributed training via Hugging Face Accelerate, and native Kubernetes support. Linux Foundation AI project. Apache 2.0…
A reliable, scalable and easy to use TensorFlow library for contextual bandits and reinforcement learning.
An open-source deep reinforcement learning framework, with an emphasis on modularized flexible library design and straightforward usability for applications in research and practice.
Neural Networks on top of tensorflow, examples. keras-contrib - Keras community contributions. keras-tuner - Hyperparameter tuning for Keras. hyperas - Keras + Hyperopt: Convenient hyperparameter optimization wrapper. elephas - Distributed Deep learning with Keras & Spark. tflearn - Neural…
Keras community contributions.
Keras + Hyperopt: A very simple wrapper for convenient hyperparameter optimization.
Distributed Deep learning with Keras & Spark.
Deep learning on graphs.
A quantization deep learning library.
Deep Reinforcement Learning for Keras.
Hyperparameter Optimization for TensorFlow, Keras and PyTorch.
Declarative statistical visualization library for Python. Can easily do many data transformation within the code to create graph
Three libraries for traditional charts, stock, and maps. Features a hand-drawn style theme option.
AnyChart Component for Ember CLI provides an easy way to use AnyChart JavaScript Charts with Ember Framework
Python library for interactive data visualization in the browser, with support for networks.
D3's simpler, easier to use cousin. Mostly predefined templates that you can just plug data in.
Allows the user to manipulate documents based on data to render charts in SVG.
Interactive line charts library that works with huge datasets.
Leading visualization and exploration software for all kinds of graphs and networks.
Resource for plotting a wide range of data (useful for visualizing survey data). Additional Information: GNU GENERAL PUBLIC LICENSE.
A charting library written in pure JavaScript, offering an easy way of adding interactive charts to your web site or web application.
is a 2D plotting library for creating static, animated, and interactive visualizations in Python. Matplotlib produces publication-quality figures in a variety of hardcopy formats and interactive environments across platforms.
Visualizer for deep learning and machine learning models (no Python code, but visualizes models from most Python Deep Learning frameworks).
Easy-to-use web service that allows for rapid creation of complex charts, from heatmaps to histograms. Upload data to create and style charts with Plotly's online spreadsheet. Fork others' plots.
SVG Data Visualization Generator - sunburst, circular dendrogram or multiple convex hull, for example. with tutorials: https://rawgraphs.io/learning
A Python visualization library based on matplotlib. It provides a high-level interface for drawing attractive statistical graphics.
Stock and financial charts.
Library for animated data visualizations and data stories.
Python package for the creation, manipulation, and study of the structure, dynamics, and functions of complex networks.
"Redash has support for querying multiple databases, including: Redshift, Google BigQuery, PostgreSQL, MySQL, Graphite, Presto, Google Spreadsheets, Cloudera Impala, Hive and custom scripts."
Easy way for everyone in your company to ask questions and learn from data. (Source Code) AGPL-3.0 Java/Docker
Debugging and visualization tool for machine learning and data science. It extensively leverages Jupyter Notebook to show real-time visualizations of data in running processes such as machine learning training.
is a popular Python framework for building ML & data science web apps for Python, R, Julia, and Jupyter.
Free online meta-analysis platform with 11 interactive D3.js statistical charts (forest plot, funnel plot, Galbraith, L'Abbé, Baujat, etc.), 5 effect size measures, AI literature screening, and publication-ready report export. github.com
Interactive notebook-based tool to visualize the forward pass of any PyTorch model.
The Data Science Lifecycle Process is a process for taking data science teams from Idea to Value repeatedly and sustainably. The process is documented in this repo
Template repository for data science lifecycle project
Synthetic tabular data generation using GANs, Diffusion Models, and LLMs with adversarial filtering and privacy metrics.
A general purpose recommender metrics library for fair evaluation.
A PyTorch based deep learning library for drug pair scoring.
Secure zero-knowledge encrypted file sharing (AES-256-GCM in-browser). No account required, MIT licensed, self-hostable, optional link expiry.
Software for corpus linguists and text/data mining enthusiasts. Build your own corpora in over 60 languages. Use over 50 tools/visualizations.
Representation learning on dynamic graphs.
A graph sampling library for NetworkX with a Scikit-Learn like API.
An unsupervised machine learning extension library for NetworkX with a Scikit-Learn like API.
All-in-one web-based IDE for machine learning and data science. The workspace is deployed as a Docker container and is preloaded with a variety of popular data science libraries (e.g., Tensorflow, PyTorch) and dev tools (e.g., Jupyter, VS Code)
A Python-powered shell that enables integration, management and orchestration of data science libraries mostly written in Python, allowing you to build pipelines, code and command-based workflows. It can also be used as a kernel for Jupyter Notebook.
Community-friendly platform supporting data scientists in creating and sharing machine learning models. Neptune facilitates teamwork, infrastructure management, models comparison and reproducibility.
Lightweight, Python library for fast and reproducible machine learning experimentation. Introduces very simple interface that enables clean machine learning pipeline design.
Curated collection of the neural networks, transformers and models that make your machine learning work faster and more effective.
easily explore, visualize, analyze, and transform data using familiar languages, such as Python and SQL, interactively.
is a personal, portable Hadoop environment that comes with a dozen interactive Hadoop tutorials.
is a free software environment for statistical computing and graphics.
is an opinionated collection of R packages designed for data science. All packages share an underlying design philosophy, grammar, and data structures.
IDE – powerful user interface for R. It’s free and open source, and works on Windows, Mac, and Linux.
Completely free enterprise-ready Python distribution for large-scale data processing, predictive analytics, and scientific computing
Pandas GUI
Free open-source SPSS alternative — menu-driven desktop statistics (t-tests, ANOVA, regression, survival analysis, ROC) with SPSS .sav import/export
Fast DataFrame library for Rust and Python, designed as a faster alternative to Pandas
free academic citation generator with a built-in reference checker that flags fabricated or hallucinated references. Searches 11+ scholarly databases (OpenAlex, PubMed, Semantic Scholar, CrossRef, SciELO), formats 40+ citation styles, and offers a public API. No sign-up; available in English,…
Machine Learning in Python
NumPy is fundamental for scientific computing with Python. It supports large, multi-dimensional arrays and matrices and includes an assortment of high-level mathematical functions to operate on these arrays.
Vaex is a Python library that allows you to visualize large datasets and calculate statistics at high speeds.
SciPy works with NumPy arrays and provides efficient routines for numerical integration and optimization.
Coursera Course
Take numerical, textual, image, GIS or other data and give it the Wolfram treatment, carrying out a full spectrum of data science analysis and visualization and automatically generate rich interactive reports—all powered by the revolutionary knowledge-based Wolfram Language.
The Kite Software Development Kit (Apache License, Version 2.0), or Kite for short, is a set of libraries, tools, examples, and documentation focused on making it easier to build systems on top of the Hadoop ecosystem.
Run, scale, share, and deploy your models — without any infrastructure or setup.
A platform for efficient, distributed, general-purpose data processing.
Apache Hama is an Apache Top-Level open source project, allowing you to do advanced analytics beyond MapReduce.
Weka is a collection of machine learning algorithms for data mining tasks.
GNU Octave is a high-level interpreted language, primarily intended for numerical computations.(Free Matlab)
Lightning-fast cluster computing
a service for exposing Apache Spark analytics jobs and machine learning models as realtime, batch or reactive web services.
A data science and engineering platform making Apache Spark more developer-friendly and cost-effective.
Deep Learning Framework
A SCIENTIFIC COMPUTING FRAMEWORK FOR LUAJIT
Intel® Nervana™ reference deep learning framework committed to best performance on all hardware.
High performance distributed data processing in NodeJS
A machine learning package built for humans.
Intel® Deep Learning Framework
An open source data visualization platform helping everyone to create simple, correct and embeddable charts. Also at github.com
TensorFlow is an Open Source Software Library for Machine Intelligence
An introductory yet powerful toolkit for natural language processing and classification
Industrial-grade speech recognition toolkit supporting 50+ languages with built-in VAD, punctuation, speaker diarization, and emotion detection. OpenAI-compatible API server included.
Free End-to-End No-Code platform for text annotation and DL model training/tuning. Out-of-the-box support for Named Entity Recognition, Classification, Relation extraction and Assertion Status Spark NLP models. Unlimited support for users, teams, projects, documents.
This module covers some basic nlp principles and implementations. The main focus is performance. When we deal with sample or training data in nlp, we quickly run out of memory. Therefore every implementation in this module is written as stream to only hold that data in memory that is currently…
high-level, high-performance dynamic programming language for technical computing
a Julia-language backend combined with the Jupyter interactive environment
Web-based notebook that enables data-driven, interactive data analytics and collaborative documents with SQL, Scala and more
An open source framework for automated feature engineering written in python
Cleansing, pre-processing, feature engineering, exploratory data analysis and easy ML with PySpark backend.
А fast and framework agnostic image augmentation library that implements a diverse set of augmentation techniques. Supports classification, segmentation, and detection out of the box. Was used to win a number of Deep Learning competitions at Kaggle, Topcoder and those that were a part of the CVPR…
An open-source data science version control system. It helps track, organize and make data science projects reproducible. In its very basic scenario it helps version control and share large data and model files.
is a workflow engine that significantly simplifies data analysis by combining in one analysis pipeline (i) feature engineering and machine learning (ii) model training and prediction (iii) table population and column evaluation.
A feature store for the management, discovery, and access of machine learning features. Feast provides a consistent view of feature data for both model training and model serving.
MLOps Tools For Managing & Orchestrating The Machine Learning LifeCycle. Reproducible and scalable machine learning workflows on Kubernetes with experiment tracking, model management, and pipeline orchestration. Apache 2.0 licensed.
Easy-to-use text annotation tool for teams with most comprehensive auto-annotation features. Supports NER, relations and document classification as well as OCR annotation for invoice labeling
Auto-Magical Experiment Manager, Version Control & DevOps for AI
Open-source data-intensive machine learning platform with a feature store. Ingest and manage features for both online (MySQL Cluster) and offline (Apache Hive) access, train and serve models at scale.
MindsDB is an Explainable AutoML framework for developers. With MindsDB you can build, train and use state of the art ML models in as simple as one line of code.
A Pytorch based framework that breaks down machine learning problems into smaller blocks that can be glued together seamlessly with an objective to build predictive models with one line of code.
An open-source Python package that extends the power of Pandas library to AWS connecting DataFrames and AWS data related services (Amazon Redshift, AWS Glue, Amazon Athena, Amazon EMR, etc).
AWS Rekognition is a service that lets developers working with Amazon Web Services add image analysis to their applications. Catalog assets, automate workflows, and extract meaning from your media and applications.
Automatically extract printed text, handwriting, and data from any document.
Spot product defects using computer vision to automate quality inspection. Identify missing product components, vehicle and structure damage, and irregularities for comprehensive quality control.
Automate code reviews and optimize application performance with ML-powered recommendations.
An open source toolkit for using continuous integration in data science projects. Automatically train and test models in production-like environments with GitHub Actions & GitLab CI, and autogenerate visual reports on pull/merge requests.
An open source Python library to painlessly transition your analytics code to distributed computing systems (Big Data)
A Python-based inferential statistics, hypothesis testing and regression framework
An open-source library for topic modeling of natural language text
Grid studio is a web-based spreadsheet application with full integration of the Python programming language.
Python Data Science Handbook: full text in Jupyter Notebooks
A data-driven framework to quantify the value of classifiers in a machine learning ensemble.
A platform built on open source tools for data, model and pipeline management.
A new kind of data science notebook. Jupyter-compatible, with real-time collaboration and running in the cloud.
An MLOps platform that handles machine orchestration, automatic reproducibility and deployment.
A Python Library for Probabalistic Programming (Bayesian Inference and Machine Learning)
Python interface to Stan (Bayesian inference and modeling)
Unsupervised learning and inference of Hidden Markov Models
ML powered analytics engine for outlier/anomaly detection and root cause analysis
Python library for anomaly detection on streaming data
A full-stack MLOps platform designed to help data scientists and machine learning practitioners around the world discover, create, and launch multi-cloud apps from their web browser.
A Python library that helps you encode your unstructured data into embeddings.
Ever been frustrated with cleaning up long, messy Jupyter notebooks? With LineaPy, an open source Python library, it takes as little as two lines of code to transform messy development code into production pipelines.
🏕️ machine learning development environment for data science and AI/ML engineering teams
A search engine 🔎 tool to discover & find a curated list of popular & new libraries, top authors, trending project kits, discussions, tutorials & learning resources
Python library for data-centric AI and automatically detecting various issues in ML datasets
AutoML to easily produce accurate predictions for image, text, tabular, time-series, and multi-modal data
Arize AI community tier observability tool for monitoring machine learning models in production and root-causing issues such as data quality and performance drift.
Aureo.io is a low-code platform that focuses on building artificial intelligence. It provides users with the capability to create pipelines, automations and integrate them with artificial intelligence models – all with their basic data.
Free cloud based entity relationship diagram (ERD) tool made for developers.
MLOps in a notebook - uncover insights, surface problems, monitor, and fine tune your models.
An MLOps platform with experiment tracking, model production management, a model registry, and full data lineage to support your ML workflow from training straight through to production.
Evaluate, test, and ship LLM applications across your dev and production lifecycles.
AI-powered collaborative environment for research. Find relevant papers, create collections to manage bibliography, and summarize content — all in one place
Workflow tool to automatically organize data visualization output
Experiment tracking, dataset versioning, and model management
Platform to programmatically author, schedule, and monitor workflows
Open-source Python framework for creating reproducible, maintainable data science code
InterpretML implements the Explainable Boosting Machine (EBM), a modern, fully interpretable machine learning model based on Generalized Additive Models (GAMs). This open-source package also provides visualization tools for EBMs, other glass-box models, and black-box explanations
Supercharged IDE for Data Science
A Python library to ease preprocessing and feature engineering for tabular machine learning
Framework-agnostic TypeScript library for generating, searching, and comparing MinHash fingerprints for fast text similarity, deduplication, and retrieval.
Popular open platform for sharing ML models, datasets, and collaborating on NLP and generative AI projects.
An open-source project that automatically maps relationship networks by parsing public data using LLMs and visualizes it as an interactive graph.
An open-source data profiler specifically focused on discovery and validation of complex patterns, such as numerical association rules, differential dependencies, denial constraints, and more.
Personal genome analysis toolkit with Python scripts analyzing raw DNA data across 17 categories (health risks, ancestry, pharmacogenomics, nutrition, psychology, and more) and generating a terminal-style single-page HTML visualization.
Fast MATLAB-syntax runtime with automatic CPU/GPU execution and fused array kernels.
A terminal UI for experimenting with custom rule engines and selective LLM analysis on real-time data streams, without worrying about streaming infra or backpressure.
Open source “failure atlas” of 16 recurring issues in LLM and RAG pipelines, with observable symptoms and suggested fixes for data science teams.
Track real-time GPU and LLM pricing across all cloud and inference providers.
An agentic LLM for autonomous data science, which can autonomously complete a wide range of data science tasks without human intervention.
Superhuman exploratory data analysis. Finds the feature interactions and subgroup effects in tabular data that LLMs and manual exploration miss — with p-values, effect sizes, and literature citations. Free for public data.
Chat with your database in natural language — no SQL needed. Get instant insights, build self-refreshing dashboards, and trigger automated workflows based on database changes.
AI crypto trading framework using LightGBM + XGBoost ensemble with 72 ML features. 70.9% walk-forward validated accuracy on out-of-sample data. Supports Bybit and Binance. MIT licensed, available on PyPI.
Open-source platform to simulate, evaluate, trace, guardrail, route, and optimize LLM and AI agent apps in one feedback loop, so agents don't just get monitored, they self-improve. Self-hostable. Apache-2.0.
Browser-based Jupyter notebook viewer and exporter that converts .ipynb notebooks to PDF, HTML, and Python without installing Python or TeX.
by Kevin Patrick Murphy
Early Access
by Jake VanderPlas
& (cheaper PDF version)
free eBook sampler
free eBook sampler
Early access
Early Access
Early access
Early access
free e-book comprehended by an online course
Early access
Free Download
Free Download
Free Download
Free Download
Convex Optimization book by Stephen Boyd - Free Download
Early Access
This book will teach you how to do data science with R. You will learn how to get your data into R, get it into the most useful structure, transform it, visualize it and model it. Exercise Solutions Authors: Garrett Grolemund and Hadley Wickham.
Early access
Early Access
Mathematical foundations by Ian Goodfellow, Yoshua Bengio, and Aaron Courville.
Early Access
Gareth James, Daniela Witten, Trevor Hastie, Robert Tibshirani.
by Trevor Hastie, Robert Tibshirani, and Jerome Friedman
Early Access
Early Access
Early Access
by David Mertz
Interactive deep learning book with code implementations
Free Download
Early Access
Early Access
Free html page
German book about applied data science
A free FreeCodeCamp book teaching the math behind AI in plain English from an engineering point of view.
A high-level guide to managing data science teams and projects.
A modern, open-access textbook on statistics with a heavy focus on data science applications.
Focuses on the "art" of data analysis, how to ask the right questions and refine them.
International Conference on Machine Learning
The Genetic and Evolutionary Computation Conference (GECCO)
an international journal devoted to applications of statistical methods at large
Like Hacker News, but for data
Data Science related publications on medium
Genetic Algorithm related Publications towards Data Science
AI industry research and analysis with papers on AI pricing, enterprise adoption, and evaluation frameworks.
Curated AI intelligence briefing from industry leaders covering models, funding, policy, and applications. 3x/week since 2017, 40K+ subscribers.
. A weekly newsletter about data-related things. Archive.
. A newsletter about data science. Archive.
. A free daily newsletter covering the most impactful developments in AI, ML, and tech. Archive.
. Practical AI engineering and generative AI explained simply: RAG, agents, and LLM application patterns for builders.
Weekly pandas exercises based on current events and real-world public data, with fully worked solutions. Issues older than two years are free, as are the first two questions + answers in current issues. Archive.
. This is the mailing list for the Research Software Engineering in the Digital Humanities (DH-RSE) working group.
Wes McKinney Archives.
Mining The Social Web.
Greg Reda Personal Blog
Recurse Center alumna
Personal Web Page
Personal Web Page
Personal Web Page
Personal Web Page
Personal Blog
Personal Blog
AllThings Data Sciene
Tech Blog on Master Data Management And Every Buzz Surrounding It
The Open Source Data Science Masters
by Peter Skomoroch. MACHINE LEARNING, DATA MINING, AND MORE
Data Science Questions and Answers from experts
a PhD student at Berkeley
a technology guy with a penchant for the web and for data, big and small
about helping professional programmers confidently apply machine learning algorithms to address complex problems.
Personal Blog
Weekly News Blog
Data Science Blog
R Bloggers
Big data
Yet Another Data Blog
Data Mining, Analytics, Big Data, Data, Science not a blog a portal
Personal Blog
is building the data scientist culture.
is some of, all of, or much more than the above and this blog explores its impact on information technology, the business world, government agencies, and our lives.
Magnus Notitia
How a Social Scientist Jumps into the World of Big Data
Thoughts on Statistical Computing and Visualization
Learning To Be A Data Scientist
The File Drawer - Chris Said's science blog
, University of Southern California, Los Angeles, USA
Visualization and Statistics
A Machine Learning Craftsmanship Blog
Handbook and recipes for data-driven solutions of real-world problems
A blog on the newly emerging data economy
A blog with resources for data science learners
A full-fledged website about data science and analytics study material.
Focused on Web Analytics.
Data science tutorials for beginners!
Blog for understanding Neural Networks!
Blog for NLP and transfer learning!
Data Science and AI notes
Data Science with Esoteric programming languages
Blog for Evolutionary Algorithms
Review and extract key concepts from academic papers
Data Science notebooks
Data science blog
ML Engineering, MLOps, and the use of ML in startups
Data science blog
ML,DL,Data Science blog
Data Science with Python
ML, DL and Data Science
ML, DL and Data Science
In-depth articles on AI, machine learning, and data science concepts with practical applications.
Educational content on software development, AI, and career growth in tech.
Mlu is developed amazon to help people in ml space you can learn everything from basics here with live diagrams
ML, DL and Data Science - with a focus on text-/data-mining
Guide to data sharing.
AI infrastructure and developer tools, with interviews from engineering leaders and technical founders.
Description: A show that digs deep into databases, big data pipelines, data governance, data collection, ETL, building effective data teams, and all of the other challenges faced by technology professionals working on data management.; Frequency: Once a week; Host: Tobias Macey; Runtime: 30-60…
Description: High level concepts in data science, and longer interview with researchers and practitioners; Host: Kyle Polich @DataSkeptic; Frequency: Once a week; Runtime: 15 - 45 mins, regularly ~25 mins
Description: Interviews with industry experts on what data science is, what problems it tries to solve and what it looks like in practice.; Host: Hugo Bowne-Anderson @hugobowne; Frequency: Once a week; Runtime: 50 - 60 mins
Description: Podcast that interviews various people in the data science industry.; Host: Chris Benson @chrisbenson, Daniel Whitenack @dwhitena; Frequency: Once a week; Runtime: Alternates between ~5 and ~60 mins
How analytics engineers build and maintain data pipelines at scale.
Data Science Education
Interesting class about neural networks available online for free by Hugo Larochelle, yet I have watched a few of those videos.
Searchable summaries and topic index for practical AI engineering talks and conference videos.
Rapid-fire, live tryouts for data scientists seeking to monetize their models as trading strategies
Big Data, Data Science, Predictive Modeling, Business Analytics, Hadoop, Decision and Operations Research.
Data scientist at Twitter
Dev, Design, Data Science @mattermark #hackerei
#datascientist @Ekimetrics. , #machinelearning #dataviz #DynamicCharts #Hadoop #R #Python #NLP #Bitcoin #dataenthousiast
Data Science Central is the industry's single resource for Big Data practitioners.
Data Science. Big Data. Data Hacks. Data Junkies. Data Startups. Open Data
Documenting my path from SQL Data Analyst pursuing an Engineering Master's Degree to Data Scientist
Mission is to help guide & advance careers in Data Science & Analytics
Tips and Tricks for Data Scientists around the world! #datascience #bigdata
DataViz, Security, Military
White House Data Chief, VP @ RelateIQ.
Data nerd, hacker, student of conflict.
Running with #BigData--enjoying a love/hate relationship with its hype. @iSchoolSU #DataScience Program Mgr.
Working @ GrubHub about data and pandas
KDnuggets President, Analytics/Big Data/Data Mining/Data Science expert, KDD & SIGKDD co-founder, was Chief Scientist at 2 startups, part-time philosopher.
Chief Scientist at RStudio, and an Adjunct Professor of Statistics at the University of Auckland, Stanford University, and Rice University.
Data Scientist
Data Scientist in Residence at @accel.
ReTweeting about data science
Scientist at Facebook and Julia developer. Author of Machine Learning for Hackers and Bandit Algorithms for Website Optimization. Tweets reflect my views only.
Principal Data Scientist @ Microsoft Data Science Team
Hacker - Pandas - Data Analyze
The Economist's Data Editor and co-author of Big Data (https://www.big-data-book.com/).
Data science instructor, and founder of Data School
Interactive data visualization and tools. Data flaneur.
DataScientist, PhD Astrophysicist, Top #BigData Influencer.
PhD Student. Programming, Mobile, Web. Artificial Intelligence, Intelligent Robotics Machine Learning, Data Mining, Natural Language Processing, Data Science.
Opinions of full-stack Python guy, author, instructor, currently playing Data Scientist. Occasional fathering, husbanding, organic gardening.
Mining the Social Web.
Data Scientist at BizQualify, Developer
Data @ Jawbone. Turned data into stories & products at LinkedIn. Text mining, applied machine learning, recommender systems. Ex-gamer, ex-machine coder; namer.
Visualization & interaction designer. Practical cyclist. Author of vis books: https://www.oreilly.com/pub/au/4419
Cloud Computing/ Big Data/ Open Data Analyst & Consultant. Writer, Speaker & Moderator. Gigaom Research Analyst.
Creating intelligent systems to automate tasks & improve decisions. Entrepreneur, ex-Principal Data Scientist @LinkedIn. Machine Learning, ProductRei, Networks
Solution Architect @ IBM, Master Data Management, Data Quality & Data Governance Blogger. Data Science, Hadoop, Big Data & Cloud.
Quora's data science topic
Tweet blog posts from the R blogosphere, data science conferences, and (!) open jobs for data scientists.
Computer scientist researching artificial intelligence. Data tinkerer. Community leader for @DataIsBeautiful. #OpenScience advocate.
Data Science geek @ UALR
Data scientist, genetic origamist, hardware aficionado
Social Scientist. Hacker. Facebook Data Science Team. Keywords: Experiments, Causal Inference, Statistics, Machine Learning, Economics.
#DataScience at Cisco
Data Scientist at BBVA Compass
Data nerd
Enjoys ABM, SNA, DM, ML, NLP, HI, Python, Java. Top percentile Kaggler/data scientist
Complex Event Processing, Big Data, Artificial Intelligence and Machine Learning. Passionate about programming and open-source.
InfoGov; Bigdata; Data as a Service; Data Science; Open, Social & Business Data Convergence
IT analyst with Ovum covering Big Data & data management with some systems engineering thrown in.
Data Scientist , Author , Entrepreneur. Co-founder @DataCommunityDC. Founder @DistrictDataLab. #DataScience #BigData #DataDC
Data Science @ PayPal. #NLP, #machinelearning; PhD, Carnegie Mellon alumni (Blog: https://allthingsds.wordpress.com )
Pandas (Python Data Analysis library).
Senior Manager - @Seagate Big Data Analytics @McKinsey Alum #BigData + #Analytics Evangelist #Hadoop, #Cloud, #Digital, & #R Enthusiast
The data news crew at @WNYC. Practicing data-driven journalism, making it visual, and showing our work.
Data science author
Data science author. Shares mostly about Julia programming
AI & Data Science Start-up Company based in England, UK
ML, DL and Data Science - with a focus on text-/data-mining
First Telegram Data Science channel. Covering all technical and popular staff about anything related to Data Science: AI, Big Data, Machine Learning, Statistics, general Math and the applications of former.
Beautiful posts on DS/ML theme with video or graphic visualization.
Daily ML news.
. A weekly newsletter about data-related things. Archive.
kaggle has built-in free jupyter notebook.; One can also connect to Google BigQuery to access big data.
Participate in data science competitions and help organizations.
Tutorial on Python time-series model deployment.
Specific datasets for aircraft and Automatic Dependent Surveillance-Broadcast (ADS-B) sources.
Curated open dataset of 100+ Chinese teas with category, origin, caffeine level, flavor notes, oxidation, and brewing parameters. Available as JSON and CSV.
Lifetime return-on-investment estimates for ~30K US bachelor's programs across 1,775 institutions, built from FREOPP, IPEDS, and BEA regional price data. 5 CSVs with data dictionary, CC BY 4.0, Zenodo DOI.
Structured dataset tracking 92 AI-attributed workforce reduction events affecting 453,748 workers across 12 countries and 11 sectors. JSON and CSV formats. CC-BY-4.0 licensed.
Public packaging product dataset generated from 1,000 exact-spec SKU records, with downloadable CSV and JSON files for ecommerce fulfillment and warehouse analysis.
320 measured PSA-style centering annotations (left/right and top/bottom border percentages, tilt) across 302 real eBay-listed Pokemon cards. CSV, CC BY 4.0, Zenodo DOI.
Median sold price by grade (raw, PSA 9, PSA 10) for 486 Pokemon cards, with sample size and confidence flag per card. CSV, CC BY 4.0, Zenodo DOI.
Weekly snapshots of public development and citation activity for open-source and research-native AI systems, content-addressed and byte-reproducible from public inputs. JSON and CSV per snapshot date, CC0, DOI 10.5281/zenodo.21076011.
The home of the U.S. Government's open data
Navigate the world of public data - Quickly search and analyze billions of public records published by governments, companies and organizations.
Amazon public datasets.
Nasdaq Data Link A premier source for financial, economic and alternative datasets.
Free AI-powered tool that scores U.S. congressional STOCK Act trade disclosures by significance. Machine-scored signals from 537 lawmakers's public trade filings.
(Storage, Lookup): Data sharing and storage
the central index for modern NLP datasets, with versioned, streamable loaders.
English dataset of Tokyo crime statistics across 5,078 neighborhoods × 7 years (36,222 records, 2018-2024), sourced from Tokyo Metropolitan Police open data. Includes interactive crime map, safety grading, and cost-of-living index. CC BY licensed.
A 30-metro composite ranking of how much of a $400K household income gets consumed by housing, taxes, childcare, healthcare, and transport. Open methodology, free, no email gate.
Open-data platform for Brazilian crime statistics. Neighborhood-level in Rio Grande do Sul (2.99M incidents across 79,024 neighborhoods, 2022–2025), municipality-level for MG and RJ, plus national PRF highway and DATASUS interpersonal-violence data. Free REST API, CSV/Parquet, daily updates, CC BY…
Filtered subset of NHTSA Fatality Analysis Reporting System covering 33,898 fatal crashes involving medium and heavy commercial trucks across all 50 US states, 2018-2024. Includes interactive Vision Zero Report Card comparing 19 cities, reproducible Python pipeline on GitHub, and HuggingFace…
Structured reference dataset of 156 peptide and peptide-adjacent compounds, each with a regulatory status bucket, category, route, half-life, molecular weight, CAS number, reference count, and PubChem/DrugBank/Wikidata IDs. CSV and JSON, no login, CC BY 4.0.
Extensive collection of datasets for practice in data analysis.
Large scale knowledge base originally stated by Metaweb. Later aquired by Google and used in Google Knowledge Graph.
Free and open access to global development data by The World Bank.
Connecting people with data for Philadelphia
Sample movie (with ratings), book and wiki datasets
contains data sets good for machine learning
by Hilary Mason
(Lookup): Weather, climate, coasts, oceans, and geophysics etc
(related: U.S. Climate Resilience Toolkit)
Datasets for Data Mining, Analytics and Knowledge Discovery.
provides a variety of data free of charge for uses that are freely available to the general public. Click on a data set below to learn more
Institute for Health Metrics and Evaluation - a catalog of health and demographic datasets from around the world and including IHME results
Collection of various open data sources.
Data from the UN
The GDELT Project monitors the world's broadcast, print, and web news from nearly every corner of every country in over 100 languages and identifies the people, locations, organizations, themes, sources, emotions, counts, quotes, images and events driving our global society every second of every…
an open source tool for running arbitrary queries against public data from the Stack Exchange network.
6 TB of Git repositories from GitHub.
Dataset Search is a search engine for datasets. Using a simple keyword search, users can discover datasets hosted in thousands of repositories across the Web.
An expanding dataset of historical job postings from Luxembourg from 2020 to today. Free with 250k+ job postings hosted on AWS Data Exchange.
Financial datasets (stock market data, financial statements, sustainability data, and more).
Daily open dataset of the cheapest new internal 3.5" SATA hard-drive price per terabyte (USD/TB) by capacity tier on Amazon US, with a historical time series. CSV, JSON and JSONL, no login, CC BY 4.0.
AI-powered multi-market stock analysis with transparent BDE scoring across 73 stocks (US/HK/A-share). EU AI Act Art.50 compliant. MIT license.
Free corpora (over 6 billion tokens) mostly German (both historically and in contemporary German).
CLARIN is a European repository for scientific datasets.
Global Biodiversity Information Facility: 2.4B+ species occurrence records. Free, open API for ecological modeling and ML research.
UN FAO statistics on food production, trade, land use, and emissions for 245+ countries. Free API and bulk download.
Free platform archiving 6B+ animal movement records from GPS and satellite telemetry. Open REST API, useful for spatiotemporal modeling and trajectory ML.
Open structured data on 1.9M+ species, including traits, classification, and media. Free API and bulk downloads for biodiversity and species-classification tasks.
The world's most comprehensive authoritative data source knowledge base. 210+ curated sources from governments, international organizations, and research institutions. MCP integration for AI agents. MIT licensed.
Python package for one-line access to 38 open research datasets from Latin America (health, neuroscience, mental health, economics). pip install latamdata-py.
Free ZIP-level environmental safety data for 42,000+ US ZIP codes: water quality, air quality, PFAS contamination, radon, lead, flood risk, and 11 more verticals. Public REST API, npm/PyPI packages, CC BY 4.0.
Real-time news corpus with structured bias features across 15+ dimensions (3.2M+ articles, 5,000+ sources), live financial market data (stocks, ETFs, crypto) with AI-generated analysis, ML options pricing with probability metrics and full Greeks, historical options chain data for quantitative…
Evidence-graded dietary-supplement dataset covering dosing, bioavailability by form, drug-nutrient interactions, NHANES deficiency prevalence, FDA FAERS adverse-event signals, and cost-per-effective-dose, with every clinical claim citing a PubMed PMID. CC BY 4.0, DOI 10.57967/hf/9356.
Synthetic benchmark for font family identification with 11,995 images of words set in 600 known fonts, annotated with word and per-letter boxes.
hesreallyhim/awesome-claude-code
A hand-picked collection of the finest of resources for the most awesome of agents, Claude Code, the undisputed champion of coding companions, from the unstoppable team…
VoltAgent/awesome-agent-skills
A curated collection of 1000+ agent skills from official dev teams and the community, compatible with Claude Code, Codex, Gemini CLI, Cursor, and more.
josephmisiti/awesome-machine-learning
A curated list of awesome Machine Learning frameworks, libraries and software.
EthicalML/awesome-production-machine-learning
A curated list of awesome open source libraries to deploy, monitor, version and scale your machine learning
analysis-tools-dev/static-analysis
⚙️ A curated list of static analysis (SAST) tools and linters for all programming languages, config files, build tools, and more. The focus is on tools which improve…
kyrolabs/awesome-langchain
😎 Awesome list of tools and projects with the awesome LangChain framework