Skip to content
87

Awesome Data Engineering

A curated list of data engineering tools for software developers

9.1k stars1,634 forks313 entriesLast push Sep 7, 2026 (22 days ago)License CC0-1.0

This page lists names, links and short descriptions. The original list on GitHub is the source and belongs to its authors.

Databases

RQLite

Replicated SQLite using the Raft consensus protocol.

In 9 listsDetails

MySQL

The world's most popular open source database.

In 9 listsDetails

TiDB

A distributed NewSQL database compatible with MySQL protocol.

In 10 listsDetails

Percona XtraBackup

A free, open source, complete online backup solution for all versions of Percona Server, MySQL® and MariaDB®.

mysql_utils

Pinterest MySQL Management Tools.

MariaDB

An enhanced, drop-in replacement for MySQL.

In 3 lists

PostgreSQL

The world's most advanced open source database.

In 9 listsDetails

Rivestack

Managed PostgreSQL with pgvector for AI workloads. HNSW indexing, sub-4ms latency, and built-in SQL editor with automatic embedding generation.

In 3 lists

Amazon RDS

Makes it easy to set up, operate, and scale a relational database in the cloud.

In 4 listsDetails

Crate.IO

Scalable SQL database with the NOSQL goodies.

In 4 listsDetails

Redis

An open source, BSD licensed, advanced key-value cache and store.

In 8 listsDetails

Riak

A distributed database designed to deliver maximum data availability by distributing data across multiple servers.

AWS DynamoDB

A fast and flexible NoSQL database service for all applications that need consistent, single-digit millisecond latency at any scale.

In 7 listsDetails

HyperDex

A scalable, searchable key-value store. Deprecated.

In 2 lists

SSDB

A high performance NoSQL database supporting many data structures, an alternative to Redis.

Kyoto Tycoon

A lightweight network server on top of the Kyoto Cabinet key-value database, built for high-performance and concurrency.

IonDB

A key-value store for microcontroller and IoT applications.

Cassandra

The right choice when you need scalability and high availability without compromising performance.

In 4 listsDetails

Cassandra Calculator

This simple form allows you to try out different values for your Apache Cassandra cluster and see what the impact is for your application.

CCM

A script to easily create and destroy an Apache Cassandra cluster on localhost.

ScyllaDB

NoSQL data store using the seastar framework, compatible with Apache Cassandra.

In 3 lists

HBase

The Hadoop database, a distributed, scalable, big data store.

In 4 listsDetails

AWS Redshift

A fast, fully managed, petabyte-scale data warehouse that makes it simple and cost-effective to analyze all your data using your existing business intelligence tools.

In 4 listsDetails

FiloDB

Distributed. Columnar. Versioned. Streaming. SQL.

Vertica

Distributed, MPP columnar database with extensive analytics SQL.

In 2 lists

ClickHouse

Distributed columnar DBMS for OLAP. SQL.

MongoDB

An open-source, document database designed for ease of development and scaling.

In 9 listsDetails

Percona Server for MongoDB

Percona Server for MongoDB® is a free, enhanced, fully compatible, open source, drop-in replacement for the MongoDB® Community Edition that includes enterprise-grade features and functionality.

MemDB

Distributed Transactional In-Memory Database (based on MongoDB).

Elasticsearch

Search & Analyze Data in Real Time.

In 11 listsDetails

Couchbase

The highest performing NoSQL distributed database.

In 5 listsDetails

RethinkDB

The open-source database for the realtime web.

In 3 lists

RavenDB

Fully Transactional NoSQL Document Database.

In 2 lists

ArcadeDB

Open-source multi-model database with native graph, document, key-value, and vector support. SQL, Cypher, and Gremlin query languages. Apache 2.0 license.

In 3 lists

Neo4j

The world's leading graph database.

In 4 listsDetails

Omnigraph

Typed graph database where agents branch and merge like Git. S3-native, Rust, traversal + vector + BM25 in one runtime.

In 5 listsDetails

OrientDB

2nd Generation Distributed Graph Database with the flexibility of Documents in one product with an Open Source commercial friendly license.

ArangoDB

A distributed free and open-source database with a flexible data model for documents, graphs, and key-values.

In 7 listsDetails

Titan

A scalable graph database optimized for storing and querying graphs containing hundreds of billions of vertices and edges distributed across a multi-machine cluster.

FlockDB

A distributed, fault-tolerant graph database by Twitter. Deprecated.

Actionbase

A database for user interactions (likes, views, follows) represented as graphs, with precomputed reads served in real-time.

In 3 lists

DAtomic

The fully transactional, cloud-ready, distributed database.

In 2 lists

Apache Geode

An open source, distributed, in-memory database for scale-out applications.

Gaffer

A large-scale graph database.

In 2 lists

InfluxDB

Scalable datastore for metrics, events, and real-time analytics.

In 8 listsDetails

OpenTSDB

A scalable, distributed Time Series Database.

In 3 lists

QuestDB

A relational column-oriented database designed for real-time analytics on time series and event data.

In 3 lists

kairosdb

Fast scalable time series database.

In 3 lists

Heroic

A scalable time series database based on Cassandra and Elasticsearch, by Spotify.

Druid

Column oriented distributed data store ideal for powering interactive applications.

Riak-TS

Riak TS is the only enterprise-grade NoSQL time series database optimized specifically for IoT and Time Series data.

Akumuli

A numeric time-series database. It can be used to capture, store and process time-series data in real-time. The word "akumuli" can be translated from esperanto as "accumulate".

In 2 lists

Rhombus

A time-series object store for Cassandra that handles all the complexity of building wide row indexes.

In 2 lists

Dalmatiner DB

Fast distributed metrics database.

In 2 lists

Blueflood

A distributed system designed to ingest and process time series data.

In 2 lists

Timely

A time series database application that provides secure access to time series data based on Accumulo and Grafana.

In 2 lists

Tarantool

An in-memory database and application server.

In 3 lists

GreenPlum

The Greenplum Database (GPDB) - An advanced, fully featured, open source data warehouse. It provides powerful and rapid analytics on petabyte scale data volumes.

cayley

An open-source graph database. Google.

In 5 listsDetails

Snappydata

OLTP + OLAP Database built on Apache Spark.

In 2 lists

TimescaleDB

Built as an extension on top of PostgreSQL, TimescaleDB is a time-series SQL database providing fast analytics, scalability, with automated data management on a proven storage engine.

In 2 lists

DuckDB

A fast in-process analytical database that has zero external dependencies, runs on Linux/macOS/Windows, offers a rich SQL dialect, and is free and extensible.

In 8 listsDetails

SlothDB

In-process analytical SQL database written in C++20. Reads Parquet, CSV, JSON, Avro, Arrow, SQLite, and Excel directly. Single binary, Python package, and 1.3 MB WASM build for the browser.

In 2 lists

chDB

Embedded ClickHouse — full ClickHouse SQL dialect, ~80 data formats, and 12+ source connectors (S3, Postgres, MongoDB, Kafka, Iceberg) in core. Python, Go, Rust, Node, Bun, Zig, and Ruby bindings.

zvec

An embedded vector database for on-device RAG and edge AI, the SQLite of vector databases.

In 7 listsDetails

ReductStore

High-performance blob and time-series storage, with edge deployment, selective replication, and efficient querying of multimodal data.

In 2 lists

Manticore Search

An open-source search database for full-text, vector, and hybrid search with real-time indexing and SQL.

In 4 listsDetails

Data Comparison

datacompy

A Python library that facilitates the comparison of two DataFrames in Pandas, Polars, Spark and more. The library goes beyond basic equality checks by providing detailed insights into discrepancies at both row and column levels.

In 4 listsDetails

dvt

Data Validation Tool compares data from source and target tables to ensure that they match. It provides column validation, row validation, schema validation, custom query validation, and ad hoc SQL exploration.

koala-diff

A high-performance Python library for comparing large datasets (CSV, Parquet) locally using Rust and Polars. It features zero-copy streaming to prevent OOM errors and generates interactive HTML data quality reports.

FutureSearch SDK

Python SDK that dispatches parallel web-research agents across table rows, synthesizing multi-agent findings into structured columns.

In 2 lists

Data Ingestion

DataSpoc Pipe

Data ingestion engine that connects 400+ Singer taps to Parquet files in cloud buckets (S3, GCS, Azure). Streaming, incremental, with auto-catalog.

Enrich.sh

Managed event ingestion service that converts JSON sent to a REST API into Hive-partitioned Parquet on Cloudflare R2, queryable from DuckDB, ClickHouse, BigQuery, Snowflake, and Python.

In 2 lists

enrich-companies

CLI tool to enrich CSV files with company data (financials, contacts, metadata) from 250M+ company records. Available on npm.

ingestr

CLI tool to copy data between databases with a single command. Supports 50+ sources including PostgreSQL, MySQL, MongoDB, Salesforce, Shopify to any data warehouse.

In 6 listsDetails

Kafka

Publish-subscribe messaging rethought as a distributed commit log.

In 3 lists

BottledWater

Change data capture from PostgreSQL into Kafka. Deprecated.

kafkat

Simplified command-line administration for Kafka brokers.

kafkacat

Generic command line non-JVM Apache Kafka producer and consumer.

pg-kafka

A PostgreSQL extension to produce messages to Apache Kafka.

librdkafka

The Apache Kafka C/C++ library.

In 2 lists

kafka-docker

Kafka in Docker.

kafka-manager

A tool for managing Apache Kafka.

kafka-node

Node.js client for Apache Kafka 0.8.

In 2 lists

Secor

Pinterest's Kafka to S3 distributed consumer.

In 2 lists

Kafka-logger

Kafka-winston logger for Node.js from Uber.

Kroxylicious

A Kafka Proxy, solving problems like encrypting your Kafka data at rest.

AWS Kinesis

A fully managed, cloud-based service for real-time data processing over large, distributed data streams.

In 3 lists

RabbitMQ

Robust messaging for applications.

In 3 lists

dlt

A fast&simple pipeline building library for Python data devs, runs in notebooks, cloud functions, airflow, etc.

In 2 lists

drt

OSS Reverse ETL CLI. Sync data from warehouses to business tools via YAML.

FluentD

An open source data collector for unified logging layer.

In 5 listsDetails

Embulk

An open source bulk data loader that helps data transfer between various databases, storages, file formats, and cloud services.

Apache Sqoop

A tool designed for efficiently transferring bulk data between Apache Hadoop and structured datastores such as relational databases.

Heka

Data Acquisition and Processing Made Easy. Deprecated.

In 3 lists

Gobblin

Universal data ingestion framework for Hadoop from LinkedIn.

Nakadi

An open source event messaging platform that provides a REST API on top of Kafka-like queues.

Pravega

Provides a new storage abstraction - a stream - for continuous and unbounded data.

Apache Pulsar

An open-source distributed pub-sub messaging system.

In 2 lists

AWS Data Wrangler

Utility belt to handle data on AWS.

In 3 lists

Airbyte

Open-source data integration for modern data teams.

In 2 lists

DBConvert Streams

Self-hosted database migration and change data capture (CDC) tool with built-in SQL IDE.

In 3 lists

Artie

Real-time data ingestion tool leveraging change data capture.

Sling

CLI data integration tool specialized in moving data between databases, as well as storage systems.

Meltano

CLI & code-first ELT.

In 3 lists

Singer SDK

The fastest way to build custom data extractors and loaders compliant with the Singer Spec.

Google Sheets ETL

Live import all your Google Sheets to your data warehouse.

CsvPath Framework

A delimited data preboarding framework that fills the gap between MFT and the data lake.

Estuary Flow

No/low-code data pipeline platform that handles both batch and real-time data ingestion.

In 3 lists

db2lake

Lightweight Node.js ETL framework for databases → data lakes/warehouses.

data-genie

High-performance, streaming-first ETL engine for Node.js and TypeScript with constant memory footprint.

Kreuzberg

Polyglot document intelligence library with a Rust core and bindings for Python, TypeScript, Go, and more. Extracts text, tables, and metadata from 62+ document formats for data pipeline ingestion.

pdfmux

Python PDF-to-Markdown orchestrator. Classifies each page and routes to the optimal backend (PyMuPDF, Docling, RapidOCR, Gemini Flash), emitting Markdown plus a per-page confidence score so ingestion pipelines can quarantine low-trust pages before feeding LLMs or retrieval.

In 2 lists

DataRaven

Managed cloud object storage transfers for ingestion workflows.

In 3 lists

Xquik

Real-time X (Twitter) data extraction platform with REST API (76 endpoints), 20 bulk extraction tools, account monitoring, HMAC-signed webhooks, and MCP server for AI agent integration.

In 4 listsDetails

Arpe.io

High-speed CLI tools for database export, import, replication and migration with parallel streaming to CSV, Parquet, JSON and cloud storage, supporting PostgreSQL, MySQL, Oracle, SQL Server and 80+ sources.

Crustdata

A real-time B2B data API for company and people intelligence, providing firmographics, headcount signals, job listings, web traffic, and funding events via REST API and webhooks for data enrichment pipelines.

crdt-merge

Conflict-free merge for DataFrames, JSON, ML models & distributed agents — powered by CRDTs.

LinkedIn Jobs Scraper

Crawlee-based actor extracting structured LinkedIn job listings at scale without API keys.

CARQ

Context-Aware RAG Processing Queue for high availability and adaptive rate-limiting.

Duckle

Local-first, open-source desktop ETL/ELT studio: drag a pipeline onto a canvas (or describe it to a built-in on-device AI assistant) and run it at native speed through DuckDB. 290+ connectors, a scheduler, and an MCP server for driving pipelines from an LLM. No cloud, no servers.

In 3 lists

Rawbbit

Open-source self-hosted game analytics pipeline. HTTP event collector with NATS JetStream buffering, raw Parquet in object storage you own, and ClickHouse for queries via Metabase, SQL, or a read-only MCP server for AI agents. Designed for teams that want to own their raw game data.

faucet-stream

Config-driven data-movement platform for Rust with pluggable source and sink connectors, running ETL, CDC, and streaming pipelines declaratively from YAML or embedded as a library.

Jitsu

An open-source Customer Data Platform. Captures event data from websites, apps, and servers and streams it into ClickHouse, Snowflake, BigQuery, Redshift, Postgres, and MySQL in real time.

Scriptella ETL

Open-source, Java-based ETL and script execution tool for transferring and transforming data between databases, files, and other sources.

In 2 lists

File System

HDFS

A distributed file system designed to run on commodity hardware.

Snakebite

A pure python HDFS client.

AWS S3

Object storage built to retrieve any amount of data from anywhere.

In 6 listsDetails

smart_open

Utils for streaming large files (S3, HDFS, gzip, bz2).

Alluxio

A memory-centric distributed storage system enabling reliable data sharing at memory-speed across cluster frameworks, such as Spark and MapReduce.

CEPH

A unified, distributed storage system designed for excellent performance, reliability, and scalability.

JuiceFS

A high-performance Cloud-Native file system driven by object storage for large-scale data storage.

In 9 listsDetails

OrangeFS

Orange File System is a branch of the Parallel Virtual File System.

SnackFS

A bite-sized, lightweight HDFS compatible file system built over Cassandra.

GlusterFS

Gluster Filesystem.

In 5 listsDetails

XtreemFS

Fault-tolerant distributed file system for all storage needs.

In 2 lists

SeaweedFS

Seaweed-FS is a simple and highly scalable distributed file system. There are two objectives: to store billions of files! to serve the files fast! Instead of supporting full POSIX file system semantics, Seaweed-FS choose to implement only a key~file mapping. Similar to the word "NoSQL", you can…

In 6 listsDetails

S3QL

A file system that stores all its data online using storage services like Google Storage, Amazon S3, or OpenStack.

LizardFS

Software Defined Storage is a distributed, parallel, scalable, fault-tolerant, Geo-Redundant and highly available file system.

In 2 lists

Serialization format

AKF

The AI native file format. Trust scores, source provenance, and compliance metadata that embed into 20+ formats (DOCX, PDF, images, code). EXIF for AI.

Apache Avro

Apache Avro™ is a data serialization system.

In 4 listsDetails

Apache Parquet

A columnar storage format available to any project in the Hadoop ecosystem, regardless of the choice of data processing framework, data model or programming language.

In 3 lists

Snappy

A fast compressor/decompressor. Used with Parquet.

In 2 lists

PigZ

A parallel implementation of gzip for modern multi-processor, multi-core machines.

Apache ORC

The smallest, fastest columnar storage for Hadoop workloads.

In 2 lists

Apache Thrift

The Apache Thrift software framework, for scalable cross-language services development.

In 4 listsDetails

ProtoBuf

Protocol Buffers - Google's data interchange format.

In 9 listsDetails

SequenceFile

A flat file consisting of binary key/value pairs. It is extensively used in MapReduce as input/output formats.

Kryo

A fast and efficient object graph serialization framework for Java.

In 4 listsDetails

PFC-JSONL

Specialized JSONL log compressor with block-level timestamp indexing and DuckDB integration. Achieves ~9% compression ratio (better than gzip) with time-range random access queries.

ParquetKit

Browser-based viewer, SQL workbench and converter for Parquet files powered by DuckDB-WASM. Fully client-side, no upload.

In 2 lists

Stream Processing

Apache Beam

A unified programming model that implements both batch and streaming data processing jobs that run on many execution engines.

In 6 listsDetails

Spark Streaming

Makes it easy to build scalable fault-tolerant streaming applications.

Apache Flink

A streaming dataflow engine that provides data distribution, communication, and fault tolerance for distributed computations over data streams.

In 6 listsDetails

Apache Storm

A free and open source distributed realtime computation system.

In 2 lists

Apache Samza

A distributed stream processing framework.

Apache NiFi

An easy to use, powerful, and reliable system to process and distribute data.

In 5 listsDetails

Apache Hudi

An open source framework for managing storage for real time processing, one of the most interesting feature is the Upsert.

In 2 lists

CocoIndex

An open source ETL framework to build fresh index for AI.

In 6 listsDetails

VoltDB

An ACID-compliant RDBMS which uses a shared nothing architecture.

In 2 lists

PipelineDB

The Streaming SQL Database.

In 2 lists

Spring Cloud Dataflow

Streaming and tasks execution between Spring Boot apps.

Bonobo

A data-processing toolkit for python 3.5+.

Robinhood's Faust

Forever scalable event processing & in-memory durable K/V store as a library with asyncio & static typing.

HStreamDB

The streaming database built for IoT data storage and real-time processing.

In 3 lists

Kuiper

An edge lightweight IoT data analytics/streaming software implemented by Golang, and it can be run at all kinds of resource-constrained edge devices.

Zilla

An API gateway built for event-driven architectures and streaming that supports standard protocols such as HTTP, SSE, gRPC, MQTT, and the native Kafka protocol.

In 4 listsDetails

SwimOS

A framework for building real-time streaming data processing applications that supports a wide range of ingestion sources.

In 2 lists

Pathway

Performant open-source Python ETL framework with Rust runtime, supporting 300+ data sources.

In 8 listsDetails

Batch Processing

Hadoop MapReduce

A software framework for easily writing applications which process vast amounts of data (multi-terabyte data-sets) - in-parallel on large clusters (thousands of nodes) - of commodity hardware in a reliable, fault-tolerant manner.

Spark

A multi-language engine for executing data engineering, data science, and machine learning on single-node machines or clusters.

In 7 listsDetails

Spark Packages

A community index of packages for Apache Spark.

Deep Spark

Connecting Apache Spark with different data stores. Deprecated.

Spark RDD API Examples

Examples by Zhen He.

Livy

The REST Spark Server.

In 2 lists

Delight

A free & cross platform monitoring tool (Spark UI / Spark History Server alternative).

In 3 lists

AWS EMR

A web service that makes it easy to quickly and cost-effectively process vast amounts of data.

In 2 lists

Data Mechanics

A cloud-based platform deployed on Kubernetes making Apache Spark more developer-friendly and cost-effective.

In 2 lists

Tez

An application framework which allows for a complex directed-acyclic-graph of tasks for processing data.

Bistro

A light-weight engine for general-purpose data processing including both batch and stream analytics. It is based on a novel unique data model, which represents data via functions and processes data via columns operations as opposed to having only set operations in conventional approaches like…

In 2 lists

Substation

A cloud native data pipeline and transformation toolkit written in Go.

In 5 listsDetails

dna-claude-analysis

Personal genome analysis toolkit with Python scripts analyzing raw DNA data across 17 categories (health risks, ancestry, pharmacogenomics, nutrition, psychology, etc.) and generating a terminal-style single-page HTML visualization.

In 7 listsDetails

H2O

Fast scalable machine learning API for smarter applications.

In 3 lists

Mahout

An environment for quickly creating scalable performant machine learning applications.

In 2 lists

Spark MLlib

Spark's scalable machine learning library consisting of common learning algorithms and utilities, including classification, regression, clustering, collaborative filtering, dimensionality reduction, as well as underlying optimization primitives.

Datatrax

Pure-Go classic machine learning toolkit and data engineering utilities. Eight algorithms with zero external dependencies.

In 2 lists

Zingg

Open source Master Data Management platform using machine learning for entity resolution at scale. Native to Databricks, Microsoft Fabric, Snowflake, AWS, and GCP. Golden records are maintained through a persistent Zingg ID across all systems and sources.

GraphLab Create

A machine learning platform that enables data scientists and app developers to easily create intelligent apps at scale.

In 2 lists

Giraph

An iterative graph processing system built for high scalability.

In 2 lists

Spark GraphX

Apache Spark's API for graphs and graph-parallel computation.

In 4 listsDetails

Presto

A distributed SQL query engine designed to query large data sets distributed over one or more heterogeneous data sources.

Hive

Data warehouse software facilitates querying and managing large datasets residing in distributed storage.

Hivemall

Scalable machine learning library for Hive/Hadoop.

PyHive

Python interface to Hive and Presto.

In 2 lists

Drill

Schema-free SQL Query Engine for Hadoop, NoSQL and Cloud Storage.

In 2 lists

Charts and Dashboards

Highcharts

A charting library written in pure JavaScript, offering an easy way of adding interactive charts to your web site or web application.

In 6 listsDetails

ZingChart

Fast JavaScript charts for any data set.

In 4 listsDetails

C3.js

D3-based reusable chart library.

In 4 listsDetails

D3.js

A JavaScript library for manipulating documents based on data.

In 9 listsDetails

D3Plus

D3's simpler, easier to use cousin. Mostly predefined templates that you can just plug data in.

In 2 lists

SmoothieCharts

A JavaScript Charting Library for Streaming Data.

PyXley

Python helpers for building dashboards using Flask and React.

Plotly

Flask, JS, and CSS boilerplate for interactive, web-based visualization apps in Python.

In 7 listsDetails

Apache Superset

A modern, enterprise-ready business intelligence web application.

In 5 listsDetails

Redash

Make Your Company Data Driven. Connect to any data source, easily visualize and share your data.

In 6 listsDetails

Metabase

The easy, open source way for everyone in your company to ask questions and learn from data.

In 8 listsDetails

stratif.io

Open-source, self-hosted, warehouse-native product analytics. Runs funnels, retention, and paths directly on DuckDB, Postgres, Snowflake, or ClickHouse.

In 2 lists

PyQtGraph

A pure-python graphics and GUI library built on PyQt4 / PySide and numpy. It is intended for use in mathematics / scientific / engineering applications.

In 2 lists

Seaborn

A Python visualization library based on matplotlib. It provides a high-level interface for drawing attractive statistical graphics.

In 6 listsDetails

QueryGPT

Natural language database query interface with automatic chart generation, supporting Chinese and English queries.

AI for Database

Agentic AI platform to connect any database (PostgreSQL, MySQL, MongoDB, etc.) and query in plain English; includes self-refreshing intelligent dashboards and action workflows triggered by data changes.

In 4 listsDetails

Dekart

Open-source SQL to map platform for BigQuery, Snowflake, and PostGIS.

LunaPad

Open-source analytics notebook for reusable SQL workflows, interactive reports, and AI-assisted data exploration.

In 2 lists

FlexViz

Open-source Python library for interactive cross-filter dashboards on large datasets. Zoom, pan, and selection are answered by lazy Polars aggregations on a stateless server instead of sending rows to the browser.

In 4 listsDetails

Workflow

Bonnard

Governed, multi-tenant MCP access to your customers' data. Turn your warehouse, dbt, or semantic layer into a secure, per-customer MCP for AI agents.

In 3 lists

Nika

Intent-as-code workflow engine for AI data pipelines: reviewable YAML DAGs statically checked (schema, permits, cost floor) before execution, with tamper-evident run traces.

In 6 listsDetails

OrionBelt Semantic Layer

Open-source semantic sidecar that compiles YAML-defined dimensions, measures, and metrics into optimized SQL across 8 engines (BigQuery, ClickHouse, Databricks, Dremio, DuckDB, MySQL, PostgreSQL, Snowflake). Unified REST, MCP, and Postgres wire protocol; one model powers AI agents, analytics, DQ…

In 2 lists

Bruin

End-to-end data pipeline tool that combines ingestion, transformation (SQL + Python), and data quality in a single CLI. Connects to BigQuery, Snowflake, PostgreSQL, Redshift, and more. Includes VS Code extension with live previews.

In 6 listsDetails

DataFlow

Open-source platform for data preparation, synthetic data generation, and AI/data pipelines. Includes reusable skills for automating workflow steps across data and AI tasks.

In 3 lists

Luigi

A Python module that helps you build complex pipelines of batch jobs.

In 15 listsDetails

CronQ

An application cron-like system. Used w/Luigi. Deprecated.

Cascading

Java based application development platform.

Airflow

A system to programmatically author, schedule, and monitor data pipelines.

In 13 listsDetails

Azkaban

A batch workflow job scheduler created at LinkedIn to run Hadoop jobs. Azkaban resolves the ordering through job dependencies and provides an easy-to-use web user interface to maintain and track your workflows.

In 3 lists

Oozie

A workflow scheduler system to manage Apache Hadoop jobs.

Pinball

DAG based workflow manager. Job flows are defined programmatically in Python. Support output passing between jobs.

In 3 lists

Dagster

An open-source Python library for building data applications.

In 12 listsDetails

Hamilton

A lightweight library to define data transformations as a directed-acyclic graph (DAG). If you like dbt for SQL transforms, you will like Hamilton for Python processing.

In 10 listsDetails

Kedro

A framework that makes it easy to build robust and scalable data pipelines by providing uniform project templates, data abstraction, configuration and pipeline assembly.

Dataform

An open-source framework and web based IDE to manage datasets and their dependencies. SQLX extends your existing SQL warehouse dialect to add features that support dependency management, testing, documentation and more.

In 2 lists

Dotflow

A lightweight Python library for building execution pipelines with retry, parallel execution, cron scheduling, and async support.

In 3 lists

Census

A reverse-ETL tool that let you sync data from your cloud data warehouse to SaaS applications like Salesforce, Marketo, HubSpot, Zendesk, etc. No engineering favors required—just SQL.

In 5 listsDetails

dbt

A command line tool that enables data analysts and engineers to transform data in their warehouses more effectively.

In 3 lists

Kestra

Scalable, event-driven, language-agnostic orchestration and scheduling platform to manage millions of workflows declaratively in code.

In 8 listsDetails

RudderStack

A warehouse-first Customer Data Platform that enables you to collect data from every application, website and SaaS platform, and then activate it in your warehouse and business tools.

In 4 listsDetails

PACE

An open source framework that allows you to enforce agreements on how data should be accessed, used, and transformed, regardless of the data platform (Snowflake, BigQuery, DataBricks, etc.)

OneQuery

Self-hosted gateway for safe, auditable queries for agents across approved data sources.

Prefect

An orchestration and observability platform. With it, developers can rapidly build and scale resilient code, and triage disruptions effortlessly.

In 3 lists

Multiwoven

The open-source reverse ETL, data activation platform for modern data teams.

In 4 listsDetails

SuprSend

Create automated workflows and logic using API's for your notification service. Add templates, batching, preferences, inapp inbox with workflows to trigger notifications directly from your data warehouse.

Mage

Open-source data pipeline tool for transforming and integrating data.

SQLMesh

An open-source data transformation framework for managing, testing, and deploying SQL and Python-based data pipelines with version control, environment isolation, and automatic dependency resolution.

Data Lake Management

lakeFS

An open source platform that delivers resilience and manageability to object-storage based data lakes.

In 10 listsDetails

Project Nessie

A Transactional Catalog for Data Lakes with Git-like semantics. Works with Apache Iceberg tables.

In 2 lists

Ilum

A modular Data Lakehouse platform that simplifies the management and monitoring of Apache Spark clusters across Kubernetes and Hadoop environments.

Gravitino

An open-source, unified metadata management for data lakes, data warehouses, and external catalogs.

FlightPath Data

FlightPath is a gateway to a data lake's bronze layer, protecting it from invalid external data file feeds as a trusted publisher.

rawquery

Managed lakehouse platform on Apache Iceberg with DuckDB query compute, S3 storage, Postgres wire protocol, and SQL transforms.

In 3 lists

ELK Elastic Logstash Kibana

docker-logstash

A highly configurable Logstash (1.4.4) - Docker image running Elasticsearch (1.7.0) - and Kibana (3.1.2).

elasticsearch-jdbc

JDBC importer for Elasticsearch.

In 3 lists

ZomboDB

PostgreSQL Extension that allows creating an index backed by Elasticsearch.

Docker

Gockerize

Package golang service into minimal Docker containers.

In 2 lists

Flocker

Easily manage Docker containers & their data.

In 2 lists

Rancher

RancherOS is a 20mb Linux distro that runs the entire OS as Docker containers.

Kontena

Application Containers for Masses.

Weave

Weaving Docker containers into applications.

In 3 lists

Zodiac

A lightweight tool for easy deployment and rollback of dockerized applications.

cAdvisor

Analyzes resource usage and performance characteristics of running containers.

In 8 listsDetails

Micro S3 persistence

Docker microservice for saving/restoring volume data to S3.

Rocker-compose

Docker composition tool with idempotency features for deploying apps composed of multiple containers. Deprecated.

Nomad

A cluster manager, designed for both long-lived services and short-lived batch processing workloads.

In 6 listsDetails

ImageLayers

Visualize Docker images and the layers that compose them.

Datasets >Realtime

DexPaprika

DEX data via SSE streaming across 36 blockchains. 36M+ pools, 33M+ tokens. Metered free tier, no API key required, data delayed up to 15s; real-time on the paid tier. Docs

In 6 listsDetails

Helium MCP

Remote MCP server for real-time financial data, 3.2M+ news articles, ML options pricing, and news bias analysis. Free, no API key. MCP

In 4 listsDetails

Twitter Realtime

The Streaming APIs give developers low latency access to Twitter's global stream of Tweet data.

Sorsa API

Real-time X (Twitter) data API providing tweets, profiles, search, communities and engagement metrics. Up to 50x cheaper than the official X API with 20 req/sec rate limit, JSON output.

Eventsim

Event data simulator. Generates a stream of pseudo-random events from a set of users, designed to simulate web traffic.

Eventum

Data generation platform for producing synthetic event streams with complex correlations.

Reddit

Real-time data is available including comments, submissions and links posted to reddit.

Datasets >Data Dumps

GitHub Archive

GitHub's public timeline since 2011, updated every hour.

In 2 lists

Common Crawl

Open source repository of web crawl data.

In 2 lists

Wikipedia

Wikipedia's complete copy of all wikis, in the form of Wikitext source and metadata embedded in XML. A number of raw database tables in SQL form are also available.

The Quiet-Broke Index

A 30-metro composite of US household cost burdens (housing, taxes, childcare, healthcare, transport) aggregated from Census ACS, BLS Consumer Expenditure Survey, and HUD Fair Market Rents. Open methodology, free, no email gate.

In 2 lists

FirstData

The world's most comprehensive authoritative data source knowledge base. 160+ curated sources from governments, international organizations, and research institutions with MCP integration.

In 2 lists

Mindweave Synthetic Business Data

42-table synthetic SME dataset with double-entry accounting, tax compliance (AU/US/UK), and temporal realism. CSV, SQL, Parquet, SQLite. Ideal for ETL pipeline testing.

LatAm Synth

Synthetic financial savings behavior generator for Latin America: users, savings goals, and transactions calibrated on 506K real records (2015–2024). Reproducible by seed, 100% synthetic.

Monitoring >Prometheus

Prometheus.io

An open-source service monitoring system and time series database.

In 13 listsDetails

HAProxy Exporter

Simple server that scrapes HAProxy stats and exports them via HTTP for Prometheus consumption.

In 3 lists

Signals CLI

Intent signal monitoring CLI. Track LinkedIn engagers, keyword posters, job changers, funding events. JSON output for data pipelines.

In 2 lists

Profiling >Data Profiler

Data Profiler

The DataProfiler is a Python library designed to make data analysis, monitoring, and sensitive data detection easy.

In 2 lists

YData Profiling

A general-purpose open-source data profiler for high-level analysis of a dataset.

In 2 lists

Desbordante

An open-source data profiler specifically focused on discovery and validation of complex patterns in data.

In 4 listsDetails

Schema

SchemaCrawler

Open-source and free relational database schema discovery and comprehension tool. Documents and diagrams relational database schemas from your Java programs, build tools and the command line. Find database design issues with lint, and write scripts against the database. Includes an MCP Server for…

Testing

Aegis DQ

Open-source agentic data quality framework with LLM-powered diagnosis, root-cause analysis, SQL auto-fix proposals, and 31 rule types — DuckDB, Postgres, BigQuery, Databricks, Athena, Snowflake.

In 2 lists

Grai

A data catalog tool that integrates into your CI system exposing downstream impact testing of data changes. These tests prevent data changes which might break data pipelines or BI dashboards from making it to production.

In 3 lists

DQOps

An open-source data quality platform for the whole data platform lifecycle from profiling new data sources to applying full automation of data quality monitoring.

DataKitchen

Open Source Data Observability for end-to-end Data Journey Observability, data profiling, anomaly detection, and auto-created data quality validation tests.

In 2 lists

GreatExpectation

Open Source data validation framework to manage data quality. Users can define and document “expectations” rules about how data should look and behave.

In 4 listsDetails

Provero

A vendor-neutral, declarative data quality engine. Define checks in YAML, run anywhere. Includes 16 built-in check types, SQL batch optimizer, anomaly detection, and data contracts.

Scherlok

Zero-config data quality CLI. Profiles every table on first run, then auto-detects anomalies (volume drops, schema drift, freshness misses, distribution shifts) on subsequent runs. No YAML, no rules to write. Works with Postgres, BigQuery, Snowflake, and dbt.

In 3 lists

RunSQL

Free online SQL playground for MySQL, PostgreSQL, and SQL Server. Create database structures, run queries, and share results instantly.

Spark Playground

Write, run, and test PySpark code on Spark Playground's online compiler. Access real-world sample datasets & solve interview questions to enhance your PySpark skills for data engineering roles.

daffy

Decorator-first DataFrame contracts/validation (columns/dtypes/constraints) at function boundaries. Supports Pandas/Polars/PyArrow/Modin.

In 2 lists

Snowflake Emulator

A Snowflake-compatible emulator for local development and testing.

In 2 lists

DataScreenIQ

Real-time data quality firewall for pipelines and APIs. Screens rows in milliseconds for schema drift, null spikes, type mismatches, and data anomalies with PASS / WARN / BLOCK decisions.

DataDriven

Interview practice with SQL query execution, Python, and data modeling exercises.

In 2 lists

Fixzi

JSON/XML validation and API contract monitoring tool for debugging and testing structured data.

dbmask

Open-source tool that scans SQL databases for sensitive columns, masks them with deterministic fakes, and validates the masked copy row by row against the original.

In 2 lists

Community >Forums

/r/dataengineering

News, tips, and background on Data Engineering.

In 2 lists

/r/etl

Subreddit focused on ETL.

AI Dev Jobs

Job board focused on AI, ML, and data engineering roles with 7,400+ listings, salary data, and a free REST API.

In 7 listsDetails

Community >Conferences

Data Council

The first technical conference that bridges the gap between data scientists, data engineers and data analysts.

Community >Podcasts

Chain of Thought

Interviews with AI and data infrastructure leaders on building production systems.

In 4 listsDetails

Data Engineering Podcast

The show about modern data infrastructure.

In 5 listsDetails

Latent Space

Technical deep dives on AI engineering, from model training to deployment.

Practical AI

Making AI practical, productive, and accessible to everyone.

Software Engineering Daily

Daily interviews about technical software topics, including data infrastructure.

In 2 lists

The Analytics Engineering Podcast

How analytics engineers build and maintain data pipelines at scale.

In 2 lists

The Data Stack Show

A show where they talk to data engineers, analysts, and data scientists about their experience around building and maintaining data infrastructure, delivering data and data products, and driving better outcomes across their businesses with data.

Community >Books

Snowflake Data Engineering

A practical introduction to data engineering on the Snowflake cloud data platform.

Best Data Science Books

This blog offers a curated list of top data science books, categorized by topics and learning stages, to aid readers in building foundational knowledge and staying updated with industry trends.

Architecting an Apache Iceberg Lakehouse

A guide to designing an Apache Iceberg lakehouse from scratch.

Learn AI Data Engineering in a Month of Lunches

A fast, friendly guide to integrating large language models into your data workflows.

See category
84

Awesome Big Data

oxnr/awesome-bigdata

A curated list of awesome big data frameworks, ressources and other awesomeness.

Active★ 15k645 entriesPushed 2 months ago
84

Awesome Public Datasets

awesomedata/awesome-public-datasets

A topic-centric list of HQ open datasets.

Fresh★ 79k3 entriesPushed today
83

Awesome Ada

ohenley/awesome-ada

A curated list of awesome resources related to the Ada and SPARK programming language

Fresh★ 869418 entriesPushed 7 days ago
80

Awesome Network Analysis

briatte/awesome-network-analysis

A curated list of awesome network analysis resources.

Fresh★ 4.1k717 entriesPushed 1 month ago
78

Awesome Public Real-Time Datasets and Sources

bytewax/awesome-public-real-time-datasets

A list of publicly available datasets with real-time data maintained by the team at bytewax.io

Active★ 2.9k91 entriesPushed 2 months ago
74

Awesome Streaming

manuzhang/awesome-streaming

a curated list of awesome streaming frameworks, applications, etc

Fresh★ 3k1 entriesPushed today