Awesome Opensource Data Engineering
An Awesome List of Open-Source Data Engineering Projects
ThisList] aims at providing an overview of https://opensource.org/licenses[open-source] projects related to data…
https://spark.apache.org/[ApacheSpark] - A unified analytics engine for large-scale data processing. Includes APIs in Scala, Java, Python (known as…
https://beam.apache.org/[ApacheBeam] - An open-source implementation of Google DataFlow. Provides capabilites of batch and streaming data processing…
https://flink.apache.org/[ApacheFlink] - Stateful computations over data streams.
https://trino.io/[Trino(formerly known as PrestoSQL)] - Distributed SQL Query Engine for Big Data.
https://superset.incubator.apache.org/[ApacheSuperset] - A modern, enterprise-ready business intelligence web application.
https://gethue.com/[HUE] - The Hadoop User Interface. Similar to Superset, but interfaces between RDBMS, Hive, Impala, HBase, Spark, HDFS &…
https://www.metabase.com/[Metabase] - An easy way for everyone in your company to ask questions and learn from data.
https://redash.io/[Redash] - All the tools to unlock your data.
https://delta.io/[DeltaLake] - Open-source storage framework that enables building a lakehouse architecture with compute engines including…
https://hudi.apache.org/[ApacheHudi] - Transactional data lake platform that brings database and data warehouse capabilities to the data lake. Hudi…
https://iceberg.apache.org/[ApacheIceberg] - High-performance format for huge analytic tables. Iceberg brings the reliability and simplicity of SQL…
https://debezium.io/[Debezium] - Change data capture for MySQL, Postgres, MongoDB, SQL Server and others.
https://github.com/zendesk/maxwell[Maxwell] - Maxwell's daemon, a MySQL-to-JSON Kafka producer.
https://calcite.apache.org/[ApacheCalcite] - SQL parser, building blocks for datastores.
http://cassandra.apache.org/[ApacheCassandra] - Open Source distributed wide column store, NoSQL database.
https://druid.apache.org/[ApacheDruid] - A high performance real-time analytics database.
https://hbase.apache.org/[ApacheHBase] - Open Source non-relational distributed database.
https://pinot.apache.org/[ApachePinot] - A realtime distributed OLAP datastore.
https://clickhouse.tech/[ClickHouse] - Open Source distributed column-oriented DBMS.
https://www.influxdata.com/[InfluxDB] - Purpose-Built Open Source Time Series Database.
https://min.io/[MinIO] - MinIO is a high performance, distributed object storage system and AWS S3 compatible.
https://www.postgresql.org/[Postgres] - The World's Most Advanced Open Source Relational Database.
https://questdb.io/[QuestDB] - Open Source Time Series Database with a focus on performance and simplicity.
https://github.com/lyft/amundsen[Amundsen] - metadata catalogue.
https://github.com/linkedin/datahub[DataHub] - A Generalized Metadata Search & Discovery Tool.
https://github.com/Netflix/metacat[Metacat] - Unified metadata exploration API service.
https://github.com/elementary-data/elementary-lineage[Elementary] - Data reliability solution, starting with plug-and-play data lineage and datasets operational status.
https://github.com/monosidev/monosi[Monosi] - Data observability & monitoring platform.
https://github.com/open-metadata/OpenMetadata[OpenMetadata] - Generalized metadata, search, and lineage tool.
https://drill.apache.org/[ApacheDrill] - Schema-free SQL Query Engine for Hadoop, NoSQL and Cloud Storage.
https://github.com/dremio/dremio-oss[Dremio] - A data lake engine. Provides an Apache Arrow-based query and acceleration engine together with the ability to…
http://teiid.io/[Teiid] - A relational abstraction of different information sources.
https://prestodb.io/[Presto] - Distributed SQL Query Engine for Big Data.
https://github.com/Alluxio/alluxio[Alluxio] - Scalable, multi-tiered distributed caching for HDFS, S3, Ceph, NFS, and related filestores. Provides integrations…
https://www.getdbt.com/[dbt] - Empowering data analysts and engineers to apply methodologies akin to those used by software engineers for…
https://avro.apache.org/[ApacheAvro] - A data serialization system.
https://parquet.apache.org/[ApacheParquet] - A columnar storage format.
https://orc.apache.org/[ApacheORC] - Another columnar storage format.
https://thrift.apache.org/[ApacheThrift] - Data type and service interface definitions and code generator.
https://arrow.apache.org/[ApacheArrow] - A cross-language development platform for in-memory data. It specifies a standardized, language-independent,…
https://capnproto.org/[Cap’nProto] - A data interchange format and capability-based RPC system.
https://google.github.io/flatbuffers/[FlatBuffers] - An efficient cross platform serialization library for C++, C#, C, Go, Java, JavaScript, Lobster, Lua, TypeScript,…
https://msgpack.org/index.html[MessagePack] - An efficient binary serialization format. It lets you exchange data among multiple languages like JSON.
https://developers.google.com/protocol-buffers[ProtocolBuffers] - Google's language-neutral, platform-neutral, extensible mechanism for serializing structured data.
https://camel.apache.org/[ApacheCamel] - Easily integrate various systems consuming or producing data.
https://kafka.apache.org/documentation/#connect[KafkaConnect] - Reusable framework to handle data int-and-out of Apache Kafka.
https://www.elastic.co/logstash[Logstash] - Open Source server-side data processing pipeline.
https://github.com/influxdata/telegraf[Telegraf] - a plugin-driven server agent writen in Go (deployed as a single binary with no external dependencies) for…
https://activemq.apache.org/[ApacheActiveMQ] - Flexible & Powerful Multi-Protocol Messaging.
https://kafka.apache.org/[ApacheKafka] - A distributed commit log with messaging capabilities.
https://pulsar.apache.org/[ApachePulsar] - A distributed pub-sub messaging system.
http://github.com/bsideup/liiklus[Liiklus] - An event gateway that provides reactive gRPC/RSocket access to Kafka-like systems.
https://nakadi.io/[Nakadi] - A distributed event bus that implements a RESTful API abstraction on top of Kafka-like queues].
https://nats.io/[NATS] - A simple, secure and high performance messaging system.
https://www.rabbitmq.com/[RabbitMQ] - A message broker.
https://github.com/wepay/waltz[Waltz] - A quorum-based distributed write-ahead log for replicating transactions.
https://zeromq.org/[ZeroMQ] - An open-source universal, high-performance messaging library.
https://cloudevents.io/[CloudEvents] - A specification for describing event data in a common way.
https://kafka.apache.org/documentation/streams/[ApacheKafka Streams] - A client library for building applications and microservices, where the input and output data are…
http://samza.apache.org/[ApacheSamza] - A distributed stream processing framework.
https://spark.apache.org/docs/latest/structured-streaming-programming-guide.html[ApacheSpark Structured Streaming] - A scalable and fault-tolerant stream processing engine built on the Spark SQL engine.
http://storm.apache.org/[ApacheStorm] - A distributed realtime computation system.
https://greatexpectations.io/[Greatexpectations] - Helps data teams eliminate pipeline debt, through data testing.
https://github.com/DataKitchen/data-observability-installer/[DataKitchenData Observability] - A full featured data quality profiling and data testing tool: it automatically generates tests…
https://prometheus.io/[prometheus] - An open-source systems monitoring and alerting toolkit.
https://grafana.com/[grafana] - An open-source analytics and monitoring platform.
https://github.com/treeverse/lakeFS/[lakeFS] - Repeatable, atomic and versioned data lake on top of object storage.
https://github.com/meirwah/awesome-workflow-engines[AwesomeWorkflow Engines] - A curated list of awesome open source workflow engines.
https://airflow.apache.org/[ApacheAirflow] - A platform created by community to programmatically author, schedule and monitor workflows.
https://nifi.apache.org/[ApacheNiFi] - Apache NiFi supports powerful and scalable directed graphs of data routing, transformation, and system…
https://github.com/knime/[KNIME] - KNIME Analytics Platform offers a WYSIWYG Editor for Spark-based workflows, with over 2000+ integrations. Offers…
https://github.com/PrefectHQ/prefect/[Prefect] - A workflow management system designed for modern infrastructure.
https://github.com/dagster-io/dagster/[Dagster] - A data orchestrator for machine learning, analytics, and ETL.
https://github.com/kestra-io/kestra[Kestra] - Open source data orchestration and scheduling platform with declarative syntax.
https://github.com/mage-ai/mage-ai[Mage] - Open source data orchestration and scheduling platform with a rich interactive UI for workflows.
https://www.dataengineeringpodcast.com/[DataEngineering Podcast]
https://softwareengineeringdaily.com/[SoftwareEngineering Daily]
https://datastackshow.com/[DataStack Show]
https://dataengweekly.substack.com/[DataEng Weekly]
https://nosql-database.org/[NOSQLDatabase Management Systems] - List of NoSQL database management systems.
https://db-engines.com/en/[DB-Engines] - Knowledge base of relational and NoSQL database management systems.
https://www.goodreads.com/list/show/146550.Data_Engineering_Group[Books] and https://www.goodreads.com/group/show/1073364-data-engineering[Book club] - Goodreads list and group about Data…
https://www.kdnuggets.com/25-free-books-to-master-sql-python-data-science-machine-learning-and-natural-language-process…Free Data Books] - Collection of 25 free e-books related to SQL, Python, Data Science, Machine Learning, and Natural…