Skip to content
84

Awesome Big Data

A curated list of awesome big data frameworks, ressources and other awesomeness.

15k stars2,583 forks645 entriesLast push Jul 31, 2026 (2 months ago)License MIT

This page lists names, links and short descriptions. The original list on GitHub is the source and belongs to its authors.

RDBMS

MySQL

The world's most popular open source database.

In 9 listsDetails

PostgreSQL

The world's most advanced open source database.

In 9 listsDetails

Oracle Database

object-relational database management system.

Teradata

high-performance MPP data warehouse platform.

Frameworks

Bistro

general-purpose data processing engine for both batch and stream analytics. It is based on a novel data model, which represents data via functions and processes data via column operations as opposed to having only set operations in conventional approaches like MapReduce or SQL.

IBM Streams

platform for distributed processing and real-time analytics. Integrates with many of the popular technologies in the Big Data ecosystem (Kafka, HDFS, Spark, etc.)

Apache Hadoop

framework for distributed processing. Integrates MapReduce (parallel processing), YARN (job scheduling) and HDFS (distributed file system).

In 2 lists

Tigon

High Throughput Real-time Stream Processing Framework.

Numaflow

Kubernetes-native stream processing platform.

In 2 lists

Pachyderm

Pachyderm is a data storage platform built on Docker and Kubernetes to provide reproducible data processing and analysis.

Polyaxon

A platform for reproducible and scalable machine learning and deep learning.

In 9 listsDetails

Smooks

An extensible Java framework for building XML and non-XML (CSV, EDI, Java, etc...) streaming applications.

In 3 lists

Distributed Programming

AddThis Hydra

distributed data processing and storage system originally developed at AddThis.

AMPLab SIMR

run Spark on Hadoop MapReduce v1.

Apache APEX

a unified, enterprise platform for big data stream and batch processing.

Apache Beam

an unified model and set of language-specific SDKs for defining and executing data processing workflows.

In 6 listsDetails

Apache Crunch

a simple Java API for tasks like joining and data aggregation that are tedious to implement on plain MapReduce.

In 3 lists

Apache DataFu

collection of user-defined functions for Hadoop and Pig developed by LinkedIn.

Apache Flink

high-performance runtime, and automatic program optimization.

In 2 lists

Apache Gearpump

real-time big data streaming engine based on Akka.

Apache Gora

framework for in-memory data model and persistence.

In 3 lists

Apache Hama

BSP (Bulk Synchronous Parallel) computing framework.

In 2 lists

Apache MapReduce

programming model for processing large data sets with a parallel, distributed algorithm on a cluster.

Apache Pig

high level language to express data analysis programs for Hadoop.

Apache REEF

retainable evaluator execution framework to simplify and unify the lower layers of big data systems.

In 2 lists

Apache S4

framework for stream processing, implementation of S4.

Apache Spark

framework for in-memory cluster computing.

In 3 lists

Apache Spark Streaming

framework for stream processing, part of Spark.

Apache Storm

framework for stream processing by Twitter also on YARN.

In 2 lists

Apache Samza

stream processing framework, based on Kafka and YARN.

In 2 lists

Apache Tez

application framework for executing a complex DAG (directed acyclic graph) of tasks, built on YARN.

In 3 lists

Apache Twill

abstraction over YARN that reduces the complexity of developing distributed applications.

Baidu Bigflow

an interface that allows for writing distributed computing programs providing lots of simple, flexible, powerful APIs to easily handle data of any scale.

Cascalog

data processing and querying library.

In 2 lists

Cheetah

High Performance, Custom Data Warehouse on Top of MapReduce.

Concurrent Cascading

framework for data management/analytics on Hadoop.

In 2 lists

Damballa Parkour

MapReduce library for Clojure.

Datasalt Pangool

alternative MapReduce paradigm.

DataTorrent StrAM

real-time engine is designed to enable distributed, asynchronous, real time in-memory big-data computations in as unblocked a way as possible, with minimal overhead and impact on performance.

Facebook Corona

Hadoop enhancement which removes single point of failure.

Facebook Peregrine

Map Reduce framework.

Facebook Scuba

distributed in-memory datastore.

Google Dataflow

create data pipelines to help them ingest, transform and analyze data.

Google MapReduce

map reduce framework.

Google MillWheel

fault tolerant stream processing framework.

IBM Streams

platform for distributed processing and real-time analytics. Integrates with many of the popular technologies in the Big Data ecosystem (Kafka, HDFS, Spark, etc.)

JAQL

declarative programming language for working with structured, semi-structured and unstructured data.

Kite

is a set of libraries, tools, examples, and documentation focused on making it easier to build systems on top of the Hadoop ecosystem.

Metamarkets Druid

framework for real-time analysis of large datasets.

In 4 listsDetails

Netflix PigPen

map-reduce for Clojure which compiles to Apache Pig.

In 3 lists

Nokia Disco

MapReduce framework developed by Nokia.

Onyx

Distributed computation for the cloud.

Pinterest Pinlater

asynchronous job execution system.

Pydoop

Python MapReduce and HDFS API for Hadoop.

Ray

A fast and simple framework for building and running distributed applications.

In 13 listsDetails

Rackerlabs Blueflood

multi-tenant distributed metric processing system

Skale

High performance distributed data processing in NodeJS.

Stratosphere

general purpose cluster computing framework.

Streamdrill

useful for counting activities of event streams over different time windows and finding the most active one.

streamsx.topology

Libraries to enable building IBM Streams application in Java, Python or Scala.

Tuktu

Easy-to-use platform for batch and streaming computation, built using Scala, Akka and Play!

Twitter Heron

Heron is a realtime, distributed, fault-tolerant stream processing engine from Twitter replacing Storm.

In 2 lists

Twitter Scalding

Scala library for Map Reduce jobs, built on Cascading.

In 2 lists

Twitter Summingbird

Streaming MapReduce with Scalding and Storm, by Twitter.

In 2 lists

Twitter TSAR

TimeSeries AggregatoR by Twitter.

Wallaroo

The ultrafast and elastic data processing engine. Big or fast data - no fuss, no Java needed.

Distributed Filesystem

Ambry

a distributed object store that supports storage of trillion of small immutable objects as well as billions of large objects.

Apache Hadoop

framework for distributed processing. Integrates MapReduce (parallel processing), YARN (job scheduling) and HDFS (distributed file system).

In 2 lists

Apache Kudu

Hadoop's storage layer to enable fast analytics on fast data.

BeeGFS

formerly FhGFS, parallel distributed file system.

Ceph Filesystem

software storage platform designed.

Disco DDFS

distributed filesystem.

Facebook Haystack

object storage system.

Google GFS

distributed filesystem.

Google Megastore

scalable, highly available storage.

GridGain

GGFS, Hadoop compliant in-memory file system.

JuiceFS

distributed POSIX file system built on object storage.

In 9 listsDetails

Lustre file system

high-performance distributed filesystem.

Microsoft Azure Data Lake Store

HDFS-compatible storage in Azure cloud

Quantcast File System QFS

open-source distributed file system.

Red Hat GlusterFS

scale-out network-attached storage file system.

Seaweed-FS

simple and highly scalable distributed file system.

In 6 listsDetails

Alluxio

reliable file sharing at memory speed across cluster frameworks.

Tahoe-LAFS

decentralized cloud storage system.

In 2 lists

Baidu File System

distributed filesystem.

Distributed Index

Pilosa

Open source distributed bitmap index that dramatically accelerates queries across multiple, massive data sets.

Document Data Model

Actian Versant

commercial object-oriented database management systems .

Crate Data

is an open source massively scalable data store. It requires zero administration.

In 4 listsDetails

Facebook Apollo

Facebook’s Paxos-like NoSQL database.

jumboDB

document oriented datastore over Hadoop.

LinkedIn Espresso

horizontally scalable document-oriented NoSQL data store.

MarkLogic

Schema-agnostic Enterprise NoSQL database technology.

Microsoft Azure DocumentDB

NoSQL cloud database service with protocol support for MongoDB

MongoDB

Document-oriented database system.

In 9 listsDetails

RavenDB

A transactional, open-source Document Database.

In 2 lists

RethinkDB

document database that supports queries like table joins and group by.

In 3 lists

Key Map Data Model

Apache Accumulo

distributed key/value store, built on Hadoop.

In 2 lists

Apache Cassandra

column-oriented distributed datastore, inspired by BigTable.

In 6 listsDetails

Apache HBase

column-oriented distributed datastore, inspired by BigTable.

In 5 listsDetails

Baidu Tera

an Internet-scale database, inspired by BigTable.

Facebook HydraBase

evolution of HBase made by Facebook.

Google BigTable

column-oriented distributed datastore.

Google Cloud Datastore

is a fully managed, schemaless database for storing non-relational data over BigTable.

Hypertable

column-oriented distributed datastore, inspired by BigTable.

InfiniDB

is accessed through a MySQL interface and use massive parallel processing to parallelize queries.

Tephra

Transactions for HBase.

Twitter Manhattan

real-time, multi-tenant distributed database for Twitter scale.

In 2 lists

ScyllaDB

column-oriented distributed datastore written in C++, totally compatible with Apache Cassandra.

Key-value Data Model

Aerospike

NoSQL flash-optimized, in-memory. Open source and "Server code in 'C' (not Java or Erlang) precisely tuned to avoid context switching and memory copies."

In 2 lists

Amazon DynamoDB

distributed key/value store, implementation of Dynamo paper.

In 7 listsDetails

Badger

a fast, simple, efficient, and persistent key-value store written natively in Go.

Bolt

an embedded key-value database for Go.

In 5 listsDetails

BTDB

Key Value Database in .Net with Object DB Layer, RPC, dynamic IL and much more

BuntDB

a fast, embeddable, in-memory key/value database for Go with custom indexing and geospatial support.

In 8 listsDetails

Edis

is a protocol-compatible Server replacement for Redis.

ElephantDB

Distributed database specialized in exporting data from Hadoop.

In 2 lists

EventStore

distributed time series database.

In 2 lists

GhostDB

a distributed, in-memory, general purpose key-value data store that delivers microsecond performance at any scale.

Graviton

a simple, fast, versioned, authenticated, embeddable key-value store database in pure Go(lang).

GridDB

suitable for sensor data stored in a timeseries.

HyperDex

a scalable, next generation key-value and document store with a wide array of features, including consistency, fault tolerance and high performance.

In 2 lists

Ignite

is an in-memory key-value data store providing full SQL-compliant data access that can optionally be backed by disk storage.

LinkedIn Krati

is a simple persistent data store with very low latency and high throughput.

Linkedin Voldemort

distributed key/value storage system.

Oracle NoSQL Database

distributed key-value database by Oracle Corporation.

Redis

in memory key value datastore.

In 8 listsDetails

Riak

a decentralized datastore.

In 2 lists

Storehaus

library to work with asynchronous key value stores, by Twitter.

SummitDB

an in-memory, NoSQL key/value database, with disk persistence and using the Raft consensus algorithm.

Tarantool

an efficient NoSQL database and a Lua application server.

In 3 lists

TiKV

a distributed key-value database powered by Rust and inspired by Google Spanner and HBase.

Tile38

a geolocation data store, spatial index, and realtime geofence, supporting a variety of object types including latitude/longitude points, bounding boxes, XYZ tiles, Geohashes, and GeoJSON

In 8 listsDetails

TreodeDB

key-value store that's replicated and sharded and provides atomic multirow writes.

Graph Data Model

Actionbase

a database for user interactions (likes, views, follows) with precomputed reads, supports HBase.

In 3 lists

AgensGraph

transactional graph database based on PostgreSQL.

ArcadeDB

multi-model database with graph, document, key-value, time-series and vector support.

In 3 lists

Apache Spark Bagel

implementation of Pregel, part of Spark.

ArangoDB

multi model distributed database.

In 7 listsDetails

DGraph

A scalable, distributed, low latency, high throughput graph database aimed at providing Google production level scale and throughput, with low enough latency to be serving real time user queries, over terabytes of structured data.

In 6 listsDetails

EliasDB

a lightweight graph based database that does not require any third-party libraries.

In 4 listsDetails

Facebook TAO

TAO is the distributed data store that is widely used at Facebook to store and serve the social graph.

GCHQ Gaffer

Gaffer by GCHQ is a framework that makes it easy to store large-scale graphs in which the nodes and edges have statistics.

In 2 lists

Google Cayley

open-source graph database.

In 5 listsDetails

Google Pregel

graph processing framework.

GraphX

resilient Distributed Graph System on Spark.

Gremlin

graph traversal Language.

In 2 lists

Infovore

RDF-centric Map/Reduce framework.

In 2 lists

JanusGraph

open-source, distributed graph database with multiple options for storage backends (Bigtable, HBase, Cassandra, etc.) and indexing backends (Elasticsearch, Solr, Lucene).

In 3 lists

Microsoft Graph Engine

a distributed in-memory data processing engine, underpinned by a strongly-typed in-memory key-value store and a general distributed computation engine.

Nebula Graph

distributed graph database for large-scale graphs with low-latency queries.

In 2 lists

Neo4j

graph database written entirely in Java.

In 4 listsDetails

OrientDB

document and graph database.

Phoebus

framework for large scale graph processing.

Titan

distributed graph database, built over Cassandra.

Columnar Databases

Columnar Storage

an explanation of what columnar storage is and when you might want it.

Actian Vector

column-oriented analytic database.

ClickHouse

an open-source column-oriented database management system that allows generating analytical data reports in real time.

EventQL

a distributed, column-oriented database built for large-scale event collection and analytics.

MonetDB

column store database.

Parquet

columnar storage format for Hadoop.

In 2 lists

Pivotal Greenplum

purpose-built, dedicated analytic data warehouse that offers a columnar engine as well as a traditional row-based one.

Vertica

is designed to manage large, fast-growing volumes of data and provide very fast query performance when used for data warehouses.

In 2 lists

SQream DB

A GPU powered big data database, designed for analytics and data warehousing, with ANSI-92 compliant SQL, suitable for data sets from 10TB to 1PB.

Google BigQuery

Google's cloud offering backed by their pioneering work on Dremel.

Amazon Redshift

Amazon's cloud offering, also based on a columnar datastore backend.

In 4 listsDetails

IndexR

an open-source columnar storage format for fast & realtime analytic with big data.

LocustDB

an experimental analytics database aiming to set a new standard for query performance on commodity hardware.

NewSQL Databases

Actian Ingres

commercially supported, open-source SQL relational database management system.

ActorDB

a distributed SQL database with the scalability of a KV store, while keeping the query capabilities of a relational database.

Amazon RedShift

data warehouse service, based on PostgreSQL.

BayesDB

statistic oriented SQL database.

Bedrock

a simple, modular, networked and distributed transaction layer built atop SQLite.

CitusDB

scales out PostgreSQL through sharding and replication.

In 2 lists

Cockroach

Scalable, Geo-Replicated, Transactional Datastore.

In 8 listsDetails

Comdb2

a clustered RDBMS built on optimistic concurrency control techniques.

In 2 lists

Datomic

distributed database designed to enable scalable, flexible and intelligent applications.

In 3 lists

FoundationDB

distributed database, inspired by F1.

Google F1

distributed SQL database built on Spanner.

Google Spanner

globally distributed semi-relational database.

H-Store

is an experimental main-memory, parallel database management system that is optimized for on-line transaction processing (OLTP) applications.

Haeinsa

linearly scalable multi-row, multi-table transaction library for HBase based on Percolator.

In 3 lists

HandlerSocket

NoSQL plugin for MySQL/MariaDB.

InfiniSQL

infinity scalable RDBMS.

KarelDB

a relational database backed by Apache Kafka.

Map-D

GPU in-memory database, big data analysis and visualization platform.

MemSQL

in memory SQL database witho optimized columnar storage on flash.

NuoDB

SQL/ACID compliant distributed database.

Oracle TimesTen in-Memory Database

in-memory, relational database management system with persistence and recoverability.

Pivotal GemFire XD

Low-latency, in-memory, distributed SQL data store. Provides SQL interface to in-memory table data, persistable in HDFS.

SAP HANA

is an in-memory, column-oriented, relational database management system.

SenseiDB

distributed, realtime, semi-structured database.

Sky

database used for flexible, high performance analysis of behavioral data.

SymmetricDS

open source software for both file and database synchronization.

TiDB

TiDB is a distributed SQL database. Inspired by the design of Google F1.

In 10 listsDetails

VoltDB

claims to be fastest in-memory database.

In 2 lists

yugabyteDB

open source, high-performance, distributed SQL database compatible with PostgreSQL.

In 3 lists

Time-Series Databases

Axibase Time Series Database

Integrated time series database on top of HBase with built-in visualization, rule-engine and SQL support.

In 2 lists

Chronix

a time series storage built to store time series highly compressed and for fast access times.

Cube

uses MongoDB to store time series data.

Heroic

is a scalable time series database based on Cassandra and Elasticsearch.

InfluxDB

a time series database with optimised IO and queries, supports pgsql and influx wire protocols.

In 7 listsDetails

QuestDB

high-performance, open-source SQL database for applications in financial services, IoT, machine learning, DevOps and observability.

In 3 lists

IronDB

scalable, general-purpose time series database.

Kairosdb

similar to OpenTSDB but allows for Cassandra.

In 3 lists

M3DB

a distributed time series database that can be used for storing realtime metrics at long retention.

Newts

a time series database based on Apache Cassandra.

TDengine

open-source time-series database with high-performance ingestion, SQL support, and IoT-oriented storage.

In 4 listsDetails

OpenTSDB

distributed time series database on top of HBase.

In 3 lists

Prometheus

a time series database and service monitoring system.

In 13 listsDetails

Beringei

Facebook's in-memory time-series database.

TrailDB

an efficient tool for storing and querying series of events.

Druid

Column oriented distributed data store ideal for powering interactive applications

In 2 lists

Riak-TS

Riak TS is the only enterprise-grade NoSQL time series database optimized specifically for IoT and Time Series data.

Akumuli

Akumuli is a numeric time-series database. It can be used to capture, store and process time-series data in real-time. The word "akumuli" can be translated from esperanto as "accumulate".

In 2 lists

Rhombus

A time-series object store for Cassandra that handles all the complexity of building wide row indexes.

In 2 lists

Dalmatiner DB

Fast distributed metrics database

In 2 lists

Blueflood

A distributed system designed to ingest and process time series data

In 2 lists

Timely

Timely is a time series database application that provides secure access to time series data based on Accumulo and Grafana.

In 2 lists

SiriDB

Highly-scalable, robust and fast, open source time series database with cluster functionality.

Thanos

Thanos is a set of components to create a highly available metric system with unlimited storage capacity using multiple (existing) Prometheus deployments.

VictoriaMetrics

fast, scalable and resource-effective open-source TSDB compatible with Prometheus. Single-node and cluster versions included

In 10 listsDetails

Lakehouse Table Formats

Apache Hudi

open data lakehouse platform and table format for high-throughput incremental data pipelines.

In 2 lists

Apache Iceberg

open table format for huge analytic datasets with schema evolution, hidden partitioning, and time travel.

Apache Paimon

lake format for building real-time lakehouse architectures with Flink and Spark.

Apache XTable

incubating Apache project for interoperability across lakehouse table formats.

Delta Lake

open-source storage framework for building lakehouse architectures on data lakes.

SQL-like processing

Actian SQL for Hadoop

high performance interactive SQL access to all Hadoop data.

Apache Doris

real-time analytical database for high-concurrency SQL analytics, search, and warehousing.

Apache Drill

framework for interactive analysis, inspired by Dremel.

Apache HCatalog

table and storage management layer for Hadoop.

Apache Hive

SQL-like data warehouse system for Hadoop.

In 2 lists

Apache Calcite

framework that allows efficient translation of queries involving heterogeneous and federated data.

In 2 lists

Apache Phoenix

SQL skin over HBase.

In 3 lists

Aster Database

SQL-like analytic processing for MapReduce.

chDB

in-process OLAP SQL engine powered by ClickHouse, callable from Python with native pandas/Arrow DataFrame interop.

In 3 lists

Cloudera Impala

framework for interactive analysis, Inspired by Dremel.

Concurrent Lingual

SQL-like query language for Cascading.

In 2 lists

Datasalt Splout SQL

full SQL query engine for big datasets.

Dremio

an open-source, SQL-like Data-as-a-Service Platform based on Apache Arrow.

DuckDB

in-process analytical SQL database for local analytics over files, data lakes, and data frames.

In 8 listsDetails

Facebook PrestoDB

distributed SQL query engine.

In 2 lists

Google BigQuery

framework for interactive analysis, implementation of Dremel.

Materialize

is a streaming database for real-time applications using SQL for queries and supporting a large fraction of PostgreSQL.

In 2 lists

Invantive SQL

SQL engine for online and on-premise use with integrated local data replication and 70+ connectors.

PipelineDB

an open-source relational database that runs SQL queries continuously on streams, incrementally storing results in tables.

In 2 lists

Pivotal HDB

SQL-like data warehouse system for Hadoop.

rawquery

managed lakehouse query service using DuckDB over Apache Iceberg tables on object storage.

In 3 lists

RainstorDB

database for storing petabyte-scale volumes of structured and semi-structured data.

Spark Catalyst

is a Query Optimization Framework for Spark and Shark.

In 8 listsDetails

SparkSQL

Manipulating Structured Data Using Spark.

Splice Machine

a full-featured SQL-on-Hadoop RDBMS with ACID transactions.

In 2 lists

StarRocks

high-performance MPP SQL engine for real-time analytics and lakehouse queries.

Stinger

interactive query for Hive.

Tajo

distributed data warehouse system on Hadoop.

In 3 lists

Trafodion

enterprise-class SQL-on-HBase solution targeting big data transactional or operational workloads.

Trino

distributed SQL query engine for querying large datasets across heterogeneous data sources.

In 4 listsDetails

Vector Databases

Chroma

open-source embedding database for AI applications.

In 3 lists

Infinity

AI-native database for hybrid vector, sparse vector, tensor, full-text, and structured search.

In 9 listsDetails

LanceDB

open-source embedded vector database built on the Lance columnar format.

Milvus

open-source vector database for scalable similarity search.

In 14 listsDetails

Qdrant

vector database and similarity search engine with REST, gRPC, and client SDKs.

In 4 listsDetails

Weaviate

open-source vector database for semantic search with structured filtering.

Zvec

open-source, in-process vector database for dense, sparse, and hybrid similarity search.

In 7 listsDetails

Data Ingestion

redpanda

A Kafka® replacement for mission critical systems; 10x faster. Written in C++.

Airbyte

open-source data movement platform for ELT pipelines and connector-based replication.

In 2 lists

Amazon Kinesis

real-time processing of streaming data at massive scale.

In 3 lists

Amazon Web Services Glue

serverless fully managed extract, transform, and load (ETL) service

In 3 lists

Apache Chukwa

data collection system.

In 2 lists

Apache Flume

service to manage large amount of log data.

In 4 listsDetails

Apache Kafka

distributed publish-subscribe messaging system.

In 6 listsDetails

Apache NiFi

Apache NiFi is an integrated data logistics platform for automating the movement of data between disparate systems.

In 5 listsDetails

Apache Pulsar

a distributed pub-sub messaging platform with a very flexible messaging model and an intuitive client API.

In 6 listsDetails

Apache SeaTunnel

high-performance, distributed data integration platform for batch and streaming synchronization.

Apache Sqoop

tool to transfer data between Hadoop and a structured datastore.

In 2 lists

Bruin

end-to-end data pipeline tool combining ingestion, transformations, and data quality checks.

In 6 listsDetails

Census

A reverse ETL product that let you sync data from your data warehouse to SaaS Applications. No engineering favors required—just SQL.

In 5 listsDetails

DataRaven

managed cloud object storage transfers for data ingestion workflows.

In 3 lists

DBConvert Streams

self-hosted CDC replication and database migration tool.

Debezium

open-source distributed platform for change data capture.

In 2 lists

Duckle

open-source visual ETL/ELT platform built on DuckDB with connectors, data quality checks, and lineage.

Embulk

open-source bulk data loader that helps data transfer between various databases, storages, file formats, and cloud services.

Estuary

SaaS platform based on Gazette with plug-and-play connectors.

In 3 lists

Facebook Scribe

streamed log data aggregator.

In 2 lists

Flink CDC

streaming data integration tool powered by Apache Flink.

Fluentd

tool to collect events and logs.

In 2 lists

Gazette

Distributed streaming infrastructure built on cloud storage which makes it easy to mix and match batch and streaming paradigms.

Google Photon

geographically distributed system for joining multiple continuously flowing streams of data in real-time with high scalability and low latency.

Graylog

log management platform for collecting, storing, searching, and alerting on machine data.

In 4 listsDetails

Heka

open source stream processing software system.

In 3 lists

Hevo

managed data pipeline platform for moving data from databases, SaaS apps, cloud storage, SDKs, and streaming services.

In 4 lists

Hightouch

reverse ETL platform for syncing warehouse data into business applications.

HIHO

framework for connecting disparate data sources with Hadoop.

ingestr

CLI tool for copying data between sources and destinations.

In 6 listsDetails

Kestrel

distributed message queue system.

LinkedIn Espresso

horizontally scalable document-oriented NoSQL data store.

LinkedIn Kamikaze

utility package for compressing sorted integer arrays.

LinkedIn White Elephant

log aggregator and dashboard.

In 2 lists

Logstash

a tool for managing events and logs.

In 8 listsDetails

Metricbeat

lightweight shipper for system and service metrics.

Netflix Suro

log agregattor like Storm and Samza based on Chukwa.

In 4 lists

Pinterest Secor

is a service implementing Kafka log persistance.

In 2 lists

Linkedin Gobblin

linkedin's universal data ingestion framework.

In 3 lists

Skizze

sketch data store to deal with all problems around counting and sketching using probabilistic data-structures.

In 2 lists

StreamSets Data Collector

continuous big data ingest infrastructure with a simple to use IDE.

Alooma

data pipeline as a service enabling moving data sources such as MySQL into data warehouses.

RudderStack

an open source customer data infrastructure (segment, mParticle alternative) written in go.

In 4 listsDetails

Zilla

An API gateway built for event-driven architectures and streaming that supports standard protocols such as HTTP, SSE, gRPC, MQTT and the native Kafka protocol.

In 4 listsDetails

Data Quality and Observability

DataKitchen Open Source Data Observability

open-source data observability for monitoring data journeys, data quality, and pipeline events.

Great Expectations

open-source framework for validating, documenting, and testing data quality.

In 4 listsDetails

OpenLineage

open standard and reference implementation for collecting lineage metadata from data pipelines.

Soda Core

open-source Python library and CLI for data quality tests.

Service Programming

Akka Toolkit

runtime for distributed, and fault tolerant event-driven applications on the JVM.

In 3 lists

Apache Avro

data serialization system.

In 2 lists

Apache Curator

Java libraries for Apache ZooKeeper.

In 3 lists

Apache Karaf

OSGi runtime that runs on top of any OSGi framework.

In 2 lists

Apache Thrift

framework to build binary protocols.

In 3 lists

Apache Zookeeper

centralized service for process management.

In 5 listsDetails

Google Chubby

a lock service for loosely-coupled distributed systems.

In 2 lists

Hydrosphere Mist

a service for exposing Apache Spark analytics jobs and machine learning models as realtime, batch or reactive web services.

In 3 lists

LinkedIn Espresso

horizontally scalable document-oriented NoSQL data store.

Mara

A lightweight opinionated ETL framework, halfway between plain scripts and Apache Airflow

In 3 lists

OpenMPI

message passing framework.

In 3 lists

Serf

decentralized solution for service discovery and orchestration.

In 3 lists

Spotify Luigi

a Python package for building complex pipelines of batch jobs. It handles dependency resolution, workflow management, visualization, handling failures, command line integration, and much more.

In 15 listsDetails

Spring XD

distributed and extensible system for data ingestion, real time analytics, batch processing, and data export.

Twitter Elephant Bird

libraries for working with LZOP-compressed data.

In 3 lists

Twitter Finagle

asynchronous network stack for the JVM.

Scheduling

Apache Airflow

a platform to programmatically author, schedule and monitor workflows.

In 2 lists

Apache Aurora

is a service scheduler that runs on top of Apache Mesos.

In 2 lists

Apache Falcon

data management framework.

In 3 lists

Apache Oozie

workflow job scheduler.

In 3 lists

Azure Data Factory

cloud-based pipeline orchestration for on-prem, cloud and HDInsight

Chronos

distributed and fault-tolerant scheduler.

Cronicle

Distributed, easy to install, NodeJS based, task scheduler

In 2 lists

Dagster

a data orchestrator for machine learning, analytics, and ETL.

In 12 listsDetails

Linkedin Azkaban

batch workflow job scheduler.

In 3 lists

Schedoscope

Scala DSL for agile scheduling of Hadoop jobs.

Sparrow

scheduling platform.

Machine Learning

Aim

open-source AI metadata tracker for experiments and training runs.

In 7 listsDetails

Azure ML Studio

Cloud-based AzureML, R, Python Machine Learning platform

brain

Neural networks in JavaScript.

In 6 listsDetails

Oryx

Lambda architecture on Apache Spark, Apache Kafka for real-time large scale machine learning.

In 5 listsDetails

Concurrent Pattern

machine learning library for Cascading.

convnetjs

Deep Learning in Javascript. Train Convolutional Neural Networks (or ordinary ones) in your browser.

In 4 listsDetails

DataVec

A vectorization and data preprocessing library for deep learning in Java and Scala. Part of the Deeplearning4j ecosystem.

Deeplearning4j

Fast, open deep learning for the JVM (Java, Scala, Clojure). A neural network configuration layer powered by a C++ library. Uses Spark and Hadoop to train nets on multiple GPUs and CPUs.

Decider

Flexible and Extensible Machine Learning in Ruby.

ENCOG

machine learning framework that supports a variety of advanced algorithms, as well as support classes to normalize and process data.

etcML

text classification with machine learning.

Etsy Conjecture

scalable Machine Learning in Scalding.

In 2 lists

Feast

A feature store for the management, discovery, and access of machine learning features. Feast provides a consistent view of feature data for both model training and model serving.

In 3 lists

GraphLab Create

A machine learning platform in Python with a broad collection of ML toolkits, data engineering, and deployment tools.

H2O

statistical, machine learning and math runtime with Hadoop. R and Python.

In 6 listsDetails

isolation-forest

distributed Spark and Scala implementation of isolation forest for unsupervised outlier detection.

In 2 lists

Karate Club

An unsupervised machine learning library for graph structured data. Python

In 7 listsDetails

Keras

An intuitive neural net API inspired by Torch that runs atop Theano and Tensorflow.

In 2 lists

Lambdo

Lambdo is a workflow engine which significantly simplifies the analysis process by unifying feature engineering and machine learning operations.

Little Ball of Fur

A subsampling library for graph structured data. Python

In 5 listsDetails

Mahout

An Apache-backed machine learning library for Hadoop.

In 2 lists

MLbase

distributed machine learning libraries for the BDAS stack.

MLPNeuralNet

Fast multilayer perceptron neural network library for iOS and Mac OS X.

In 3 lists

ML Workspace

All-in-one web-based IDE specialized for machine learning and data science.

In 10 listsDetails

MOA

MOA performs big data stream mining in real time, and large scale machine learning.

MonkeyLearn

Text mining made easy. Extract and classify data from text.

ND4J

A matrix library for the JVM. Numpy for Java.

Neptune

experiment tracking and model registry for research and production machine learning teams.

In 6 listsDetails

nupic

Numenta Platform for Intelligent Computing: a brain-inspired machine intelligence platform, and biologically accurate neural network based on cortical learning algorithms.

In 2 lists

PredictionIO

machine learning server built on Hadoop, Mahout and Cascading.

PyTorch Geometric Temporal

a temporal extension library for PyTorch Geometric .

In 5 listsDetails

RL4J

Reinforcement learning for Java and Scala. Includes Deep-Q learning and A3C algorithms, and integrates with Open AI's Gym. Runs in the Deeplearning4j ecosystem.

SAMOA

distributed streaming machine learning framework.

scikit-learn

scikit-learn: machine learning in Python.

In 10 listsDetails

Shapley

A data-driven framework to quantify the value of classifiers in a machine learning ensemble.

In 4 listsDetails

Spark MLlib

a Spark implementation of some common machine learning (ML) functionality.

Sibyl

System for Large Scale Machine Learning at Google.

TensorFlow

Library from Google for machine learning using data flow graphs.

In 23 listsDetails

Theano

A Python-focused machine learning library supported by the University of Montreal.

Torch

A deep learning library with a Lua API, supported by NYU and Facebook.

Velox

System for serving machine learning predictions.

Vowpal Wabbit

learning system sponsored by Microsoft and Yahoo!.

WEKA

suite of machine learning software.

In 3 lists

BidMach

CPU and GPU-accelerated Machine Learning Library.

In 2 lists

Benchmarking

Apache JMeter

load testing tool for measuring performance of services and distributed systems.

In 4 lists

Berkeley SWIM Benchmark

real-world big data workload benchmark.

Estuary Benchmark Report

reproducible, vendor-neutral data warehouse benchmark.

Intel HiBench

a Hadoop benchmark suite.

In 2 lists

PUMA Benchmarking

benchmark suite for MapReduce applications.

Yahoo Gridmix3

Hadoop cluster benchmarking from Yahoo engineer team.

Deeplearning4j Benchmarks

UCSB

extended Yahoo Cloud Serving Benchmark for NoSQL databases.

Security

Apache Ranger

Central security admin & fine-grained authorization for Hadoop

Apache Eagle

real time monitoring solution

Apache Knox Gateway

single point of secure access for Hadoop clusters.

Apache Sentry

security module for data stored in Hadoop.

BDA

The vulnerability detector for Hadoop and Spark

FileShot

zero-knowledge encrypted file transfer for sharing large datasets.

In 3 lists

System Deployment

Apache Ambari

operational framework for Hadoop management.

In 3 lists

Apache Bigtop

system deployment framework for the Hadoop ecosystem.

In 3 lists

Apache Helix

cluster management framework.

In 3 lists

Apache Mesos

cluster manager.

In 5 listsDetails

Apache Slider

is a YARN application to deploy existing distributed applications on YARN.

Apache Whirr

set of libraries for running cloud services.

Apache YARN

Cluster manager.

Brooklyn

library that simplifies application deployment and management.

Buildoop

Similar to Apache BigTop based on Groovy language.

Cloudera HUE

web application for interacting with Hadoop.

In 3 lists

Facebook Prism

multi datacenters replication system.

Google Borg

job scheduling and monitoring system.

Google Omega

job scheduling and monitoring system.

Hortonworks HOYA

application that can deploy HBase cluster on YARN.

Kubernetes

a system for automating deployment, scaling, and management of containerized applications.

In 10 listsDetails

Marathon

Mesos framework for long-running services.

In 2 lists

Linkis

Linkis helps easily connect to various back-end computation/storage engines.

Terraform

infrastructure as code tool for provisioning and managing cloud and on-premises infrastructure.

In 6 listsDetails

Applications

411

an web application for alert management resulting from scheduled searches into Elasticsearch.

In 2 lists

Adobe spindle

Next-generation web analytics processing with Scala, Spark, and Parquet.

Apache Metron

a platform that integrates a variety of open source big data technologies in order to offer a centralized tool for security monitoring and analysis.

Apache Nutch

open source web crawler.

In 4 lists

Apache OODT

capturing, processing and sharing of data for NASA's scientific archives.

In 2 lists

Apache Tika

content analysis toolkit.

Argus

Time series monitoring and alerting platform.

AthenaX

a streaming analytics platform that enables users to run production-quality, large scale streaming analytics using Structured Query Language (SQL).

Atlas

a backend for managing dimensional time series data.

In 3 lists

Countly

open source mobile and web analytics platform, based on Node.js & MongoDB.

In 6 listsDetails

Comet

Comet provides an end-to-end model evaluation platform for AI developers, with best in class LLM evaluations, experiment tracking, and production monitoring.

Domino

Run, scale, share, and deploy models — without any infrastructure.

In 4 listsDetails

Eclipse BIRT

Eclipse-based reporting system.

ElastAert

ElastAlert is a simple framework for alerting on anomalies, spikes, or other patterns of interest from data in ElasticSearch.

In 4 listsDetails

Eventhub

open source event analytics platform.

In 2 lists

Gigasheet

cloud spreadsheet for exploring and analyzing large datasets.

HASH

open source simulation and visualization platform.

In 3 lists

Hermes

asynchronous message broker built on top of Kafka.

In 2 lists

Hunk

Splunk analytics for Hadoop.

Indicative

Web & mobile analytics tool, with data warehouse (AWS, BigQuery) integration.

In 2 lists

Jupyter

Notebook and project application for interactive data science and scientific computing across all programming languages.

In 6 listsDetails

MADlib

data-processing library of an RDBMS to analyze data.

Kapacitor

an open source framework for processing, monitoring, and alerting on time series data.

Kylin

open source Distributed Analytics Engine from eBay.

In 2 lists

PivotalR

R on Pivotal HD / HAWQ and PostgreSQL.

Opik

Debug, evaluate, and monitor your LLM applications, RAG systems, and agentic workflows with comprehensive tracing, automated evaluations, and production-ready dashboards.

In 6 listsDetails

Rakam

open-source real-time custom analytics platform powered by Postgresql, Kinesis and PrestoDB.

Qubole

auto-scaling Hadoop cluster, built-in data connectors.

In 2 lists

SnappyData

a distributed in-memory data store for real-time operational analytics, delivering stream analytics, OLTP (online transaction processing) and OLAP (online analytical processing) built on Spark in a single integrated cluster.

In 2 lists

Snowplow

enterprise-strength web and event analytics, powered by Hadoop, Kinesis, Redshift and Postgres.

In 2 lists

SparkR

R frontend for Spark.

Splunk

analyzer for machine-generated data.

In 4 lists

Sumo Logic

cloud based analyzer for machine-generated data.

In 2 lists

Substation

Substation is a cloud native data pipeline and transformation toolkit written in Go.

In 5 listsDetails

Talend

unified open source environment for YARN, Hadoop, HBASE, Hive, HCatalog & Pig.

Search engine and framework

Apache Lucene

Search engine library.

Apache Solr

Search platform for Apache Lucene.

In 2 lists

Elassandra

is a fork of Elasticsearch modified to run on top of Apache Cassandra in a scalable and resilient peer-to-peer architecture.

ElasticSearch

Search and analytics engine based on Apache Lucene.

In 11 listsDetails

Enigma.io

Freemium robust web application for exploring, filtering, analyzing, searching and exporting massive datasets scraped from across the Web.

In 2 lists

Google Caffeine

continuous indexing system.

Google Percolator

continuous indexing system.

HBase Coprocessor

implementation of Percolator, part of HBase.

Lily HBase Indexer

quickly and easily search for any content stored in HBase.

In 2 lists

LinkedIn Bobo

is a Faceted Search implementation written purely in Java, an extension to Apache Lucene.

LinkedIn Cleo

is a flexible software library for enabling rapid development of partial, out-of-order and real-time typeahead search.

In 2 lists

LinkedIn Galene

search architecture at LinkedIn.

In 2 lists

LinkedIn Zoie

is a realtime search/indexing system written in Java.

MG4J

MG4J (Managing Gigabytes for Java) is a full-text search engine for large document collections written in Java. It is highly customisable, high-performance and provides state-of-the-art features and new research algorithms.

Sphinx Search Server

fulltext search engine.

Vespa

is an engine for low-latency computation over large data sets. It stores and indexes your data such that queries, selection and processing over the data can be performed at serving time.

Facebook Faiss

is a library for efficient similarity search and clustering of dense vectors. It contains algorithms that search in sets of vectors of any size, up to ones that possibly do not fit in RAM. It also contains supporting code for evaluation and parameter tuning. Faiss is written in C++ with complete…

In 8 listsDetails

Annoy

is a C++ library with Python bindings to search for points in space that are close to a given query point. It also creates large read-only file-based data structures that are mmapped into memory so that many processes may share the same data.

In 6 listsDetails

MySQL forks and evolutions

Amazon RDS

MySQL databases in Amazon's cloud.

In 4 listsDetails

Drizzle

evolution of MySQL 6.0.

Google Cloud SQL

MySQL databases in Google's cloud.

MariaDB

enhanced, drop-in replacement for MySQL.

In 3 lists

MySQL Cluster

MySQL implementation using NDB Cluster storage engine.

Percona Server

enhanced, drop-in replacement for MySQL.

ProxySQL

High Performance Proxy for MySQL.

TokuDB

TokuDB is a storage engine for MySQL and MariaDB.

WebScaleSQL

is a collaboration among engineers from several companies that face similar challenges in running MySQL at scale.

PostgreSQL forks and evolutions

HadoopDB

hybrid of MapReduce and DBMS.

IBM Netezza

high-performance data warehouse appliances.

Postgres-XL

Scalable Open Source PostgreSQL-based Database Cluster.

RecDB

Open Source Recommendation Engine Built Entirely Inside PostgreSQL.

Stado

open source MPP database system solely targeted at data warehousing and data mart applications.

Yahoo Everest

multi-peta-byte database / MPP derived by PostgreSQL.

TimescaleDB

An open-source time-series database optimized for fast ingest and complex queries

PipelineDB

an open-source relational database that runs SQL queries continuously on streams, incrementally storing results in tables.

In 2 lists

Memcached forks and evolutions

Facebook McDipper

key/value cache for flash storage.

Facebook Memcached

fork of Memcache.

Twemproxy

A fast, light-weight proxy for memcached and redis.

In 2 lists

Twitter Fatcache

key/value cache for flash storage.

Twitter Twemcache

fork of Memcache.

Embedded Databases

Actian Ingres

commercially supported, open-source SQL relational database management system.

BerkeleyDB

a software library that provides a high-performance embedded database for key/value data.

HanoiDB

Erlang LSM BTree Storage.

LevelDB

a fast key-value storage library written at Google that provides an ordered mapping from string keys to string values.

In 6 listsDetails

LMDB

ultra-fast, ultra-compact key-value embedded data store developed by Symas.

RocksDB

embeddable persistent key-value store for fast storage based on LevelDB.

Business Intelligence

BIME Analytics

business intelligence platform in the cloud.

Blazer

business intelligence made simple.

In 3 lists

Chartio

lean business intelligence platform to visualize and explore your data.

In 3 lists

Count

notebook-based anlytics and visualisation platform using SQL or drag-and-drop.

In 4 lists

datapine

self-service business intelligence tool in the cloud.

In 2 lists

Dekart

Large scale geospatial analytics for Google BigQuery based on Kepler.gl.

In 2 lists

GoodData

platform for data products and embedded analytics.

In 2 lists

Jaspersoft

powerful business intelligence suite.

Jedox Palo

customisable Business Intelligence platform.

Jethrodata

Interactive Big Data Analytics.

intermix.io

Performance Monitoring for Amazon Redshift

In 2 lists

Lightdash

The open source Looker alternative built on dbt

In 2 lists

Metabase

The simplest, fastest way to get business intelligence and analytics to everyone in your company.

In 8 listsDetails

Microsoft

business intelligence software and platform.

Microstrategy

software platforms for business intelligence, mobile intelligence, and network applications.

In 2 lists

Numeracy

Fast, clean SQL client and business intelligence.

Pentaho

business intelligence platform.

In 2 lists

Qlik

business intelligence and analytics platform.

Query.me

collaborative SQL notebooks for querying, scheduling, and sharing reporting workflows.

In 3 lists

Redash

Open source business intelligence platform, supporting multiple data sources and planned queries.

In 6 listsDetails

Datapallas

BI and data platform with AI exploration, dashboards, and pixel-perfect report generation; formerly ReportBurster.

Saiku Analytics

Open source analytics platform.

In 2 lists

Knowage

open source business intelligence platform. (former SpagoBi)

SparklineData SNAP

modern B.I platform powered by Apache Spark.

Tableau

business intelligence platform.

In 7 listsDetails

Zoomdata

Big Data Analytics.

Data Visualization

Airpal

Web UI for PrestoDB.

In 2 lists

AnyChart

fast, simple and flexible JavaScript (HTML5) charting library featuring pure JS API.

In 2 lists

Arbor

graph visualization library using web workers and jQuery.

In 3 lists

Banana

visualize logs and time-stamped data stored in Solr. Port of Kibana.

In 2 lists

Bloomery

Web UI for Impala.

Bokeh

A powerful Python interactive visualization library that targets modern web browsers for presentation, with the goal of providing elegant, concise construction of novel graphics in the style of D3.js, but also delivering this capability with high-performance interactivity over very large or…

C3

D3-based reusable chart library

In 2 lists

CartoDB

open-source or freemium hosting for geospatial databases with powerful front-end editing capabilities and a robust API.

In 2 lists

chartd

responsive, retina-compatible charts with just an img tag.

In 2 lists

Chart.js

open source HTML5 Charts visualizations.

In 2 lists

Chartist.js

another open source HTML5 Charts visualization.

In 4 listsDetails

Crossfilter

JavaScript library for exploring large multivariate datasets in the browser. Works well with dc.js and d3.js.

Cubism

JavaScript library for time series visualization.

In 4 listsDetails

Cytoscape

JavaScript library for visualizing complex networks.

DC.js

Dimensional charting built to work natively with crossfilter rendered using d3.js. Excellent for connecting charts/additional metadata to hover events in D3.

D3

javaScript library for manipulating documents.

In 9 listsDetails

D3.compose

Compose complex, data-driven visualizations from reusable charts and components.

D3Plus

A fairly robust set of reusable charts and styles for d3.js.

Dash

Analytical Web Apps for Python, R, Julia, and Jupyter. Built on top of plotly, no JS required

In 7 listsDetails

Dekart

Large scale geospatial analytics for Google BigQuery based on Kepler.gl.

In 2 lists

DevExtreme React Chart

High-performance plugin-based React chart for Bootstrap and Material Design.

In 4 lists

Echarts

Baidus enterprise charts.

In 6 listsDetails

Envisionjs

dynamic HTML5 visualization.

In 2 lists

Flexmonster Pivot Table & Charts

JavaScript component for pivot tables, charts, and web reporting.

FnordMetric

write SQL queries that return SVG charts rather than tables

Frappe Charts

GitHub-inspired simple and modern SVG charts for the web with zero dependencies.

Freeboard

pen source real-time dashboard builder for IOT and other web mashups.

In 7 listsDetails

Gephi

An award-winning open-source platform for visualizing and manipulating large graphs and network connections. It's like Photoshop, but for graphs. Available for Windows and Mac OS X.

In 8 listsDetails

Google Charts

simple charting API.

In 4 listsDetails

Grafana

graphite dashboard frontend, editor and graph composer.

In 11 listsDetails

Graphite

scalable Realtime Graphing.

Highcharts

simple and flexible charting API.

In 6 listsDetails

IPython

provides a rich architecture for interactive computing.

In 2 lists

Kibana

visualize logs and time-stamped data

In 7 listsDetails

Lumify

open source big data analysis and visualization platform

Matplotlib

plotting with Python.

In 7 listsDetails

Metricsgraphic.js

a library built on top of D3 that is optimized for time-series data

In 2 lists

NVD3

chart components for d3.js.

In 2 lists

Peity

Progressive SVG bar, line and pie charts.

In 3 lists

Plot.ly

Easy-to-use web service that allows for rapid creation of complex charts, from heatmaps to histograms. Upload data to create and style charts with Plotly's online spreadsheet. Fork others' plots.

In 4 listsDetails

Plotly.js

The open source javascript graphing library that powers plotly.

In 6 listsDetails

Recline

simple but powerful library for building data applications in pure Javascript and HTML.

In 2 lists

Redash

open-source platform to query and visualize data.

In 11 listsDetails

ReCharts

A composable charting library built on React components

In 3 lists

Shiny

a web application framework for R.

Sigma.js

JavaScript library dedicated to graph drawing.

In 4 listsDetails

Superset

a data exploration platform designed to be visual, intuitive and interactive, making it easy to slice, dice and visualize data and perform analytics at the speed of thought.

In 5 listsDetails

Vega

a visualization grammar.

In 3 lists

WebDataRocks

free web pivot table component for embedding analytics in applications.

Zeppelin

a notebook-style collaborative data analysis.

Zing Charts

JavaScript charting library for big data.

In 4 listsDetails

DataSphere Studio

one-stop data application development management portal.

Internet of things and sensor data

Apache Edgent (Incubating)

a programming model and micro-kernel style runtime that can be embedded in gateways and small footprint edge devices enabling local, real-time, analytics on the edge devices.

Azure IoT Hub

Cloud-based bi-directional monitoring and messaging hub

In 3 lists

TempoIQ

Cloud-based sensor analytics.

2lemetry

Platform for Internet of things.

Pubnub

Data stream network

ThingWorx

Rapid development and connection of intelligent systems

IFTTT

If this then that

In 8 listsDetails

Evrything

Making products smart

NetLytics

Analytics platform to process network data on Spark.

Ably

Pub/sub messaging platform for IoT

In 2 lists

Interesting Readings

Big Data Benchmark

Benchmark of Redshift, Hive, Shark, Impala and Stiger/Tez.

In 2 lists

NoSQL Comparison

Cassandra vs MongoDB vs CouchDB vs Redis vs Riak vs HBase vs Couchbase vs Neo4j vs Hypertable vs ElasticSearch vs Accumulo vs VoltDB vs Scalaris comparison.

Monitoring Kafka performance

Guide to monitoring Apache Kafka, including native methods for metrics collection.

Monitoring Hadoop performance

Guide to monitoring Hadoop, with an overview of Hadoop architecture, and native methods for metrics collection.

In 2 lists

Monitoring Cassandra performance

Guide to monitoring Cassandra, including native methods for metrics collection.

Interesting Papers >2015 - 2016

2015

Facebook - One Trillion Edges: Graph Processing at Facebook-Scale.

Interesting Papers >2013 - 2014

2014

Stanford - Mining of Massive Datasets.

2013

AMPLab - Presto: Distributed Machine Learning and Graph Processing with Sparse Matrices.

2013

AMPLab - MLbase: A Distributed Machine-learning System.

2013

AMPLab - Shark: SQL and Rich Analytics at Scale.

2013

AMPLab - GraphX: A Resilient Distributed Graph System on Spark.

2013

Google - HyperLogLog in Practice: Algorithmic Engineering of a State of The Art Cardinality Estimation Algorithm.

2013

Microsoft - Scalable Progressive Analytics on Big Data in the Cloud.

2013

Metamarkets - Druid: A Real-time Analytical Data Store.

2013

Google - Online, Asynchronous Schema Change in F1.

2013

Google - F1: A Distributed SQL Database That Scales.

2013

Google - MillWheel: Fault-Tolerant Stream Processing at Internet Scale.

2013

Facebook - Scuba: Diving into Data at Facebook.

2013

Facebook - Unicorn: A System for Searching the Social Graph.

2013

Facebook - Scaling Memcache at Facebook.

Interesting Papers >2011 - 2012

2012

Twitter - The Unified Logging Infrastructure for Data Analytics at Twitter.

In 2 lists

2012

AMPLab - Blink and It’s Done: Interactive Queries on Very Large Data.

2012

AMPLab - Fast and Interactive Analytics over Hadoop Data with Spark.

2012

AMPLab - Shark: Fast Data Analysis Using Coarse-grained Distributed Memory.

2012

Microsoft - Paxos Replicated State Machines as the Basis of a High-Performance Data Store.

2012

Microsoft - Paxos Made Parallel.

2012

AMPLab - BlinkDB: Queries with Bounded Errors and Bounded Response Times on Very Large Data.

2012

Google - Processing a trillion cells per mouse click.

2012

Google - Spanner: Google’s Globally-Distributed Database.

2011

AMPLab - Scarlett: Coping with Skewed Popularity Content in MapReduce Clusters.

2011

AMPLab - Mesos: A Platform for Fine-Grained Resource Sharing in the Data Center.

2011

Google - Megastore: Providing Scalable, Highly Available Storage for Interactive Services.

Interesting Papers >2001 - 2010

2010

Facebook - Finding a needle in Haystack: Facebook’s photo storage.

2010

AMPLab - Spark: Cluster Computing with Working Sets.

Google Pregel

graph processing framework.

2010

Google - Large-scale Incremental Processing Using Distributed Transactions and notifications base of Percolator and Caffeine.

2010

Google - Dremel: Interactive Analysis of Web-Scale Datasets.

2010

Yahoo - S4: Distributed Stream Computing Platform.

2009

HadoopDB: An Architectural Hybrid of MapReduce and DBMS Technologies for Analytical Workloads.

2008

AMPLab - Chukwa: A large-scale monitoring system.

2007

Amazon - Dynamo: Amazon’s Highly Available Key-value Store.

In 2 lists

2006

Google - The Chubby lock service for loosely-coupled distributed systems.

2006

Google - Bigtable: A Distributed Storage System for Structured Data.

2004

Google - MapReduce: Simplied Data Processing on Large Clusters.

Google GFS

distributed filesystem.

Videos

Spark in Motion

Spark in Motion teaches you how to use Spark for batch and streaming data analytics.

Machine Learning, Data Science and Deep Learning with Python

LiveVideo tutorial that covers machine learning, Tensorflow, artificial intelligence, and neural networks.

In 3 lists

Data warehouse schema design - dimensional modeling and star schema

Introduction to schema design for data warehouse using the star schema method.

Elasticsearch 7 and Elastic Stack

LiveVideo tutorial that covers searching, analyzing, and visualizing big data on a cluster with Elasticsearch, Logstash, Beats, Kibana, and more.

In 2 lists

Books

Data Science at Scale with Python and Dask

Data Science at Scale with Python and Dask teaches you how to build distributed data projects that can handle huge amounts of data.

Streaming Data

Streaming Data introduces the concepts and requirements of streaming and real-time data systems.

Storm Applied

Storm Applied is a practical guide to using Apache Storm for the real-world tasks associated with processing and analyzing real-time data streams.

Fundamentals of Stream Processing: Application Design, Systems, and Analytics

This comprehensive, hands-on guide combining the fundamental building blocks and emerging research in stream processing is ideal for application designers, system builders, analytic developers, as well as students and researchers in the field.

Stream Data Processing: A Quality of Service Perspective

Presents a new paradigm suitable for stream and complex event processing.

Unified Log Processing

Unified Log Processing is a practical guide to implementing a unified log of event streams (Kafka or Kinesis) in your business

Kafka Streams in Action

Kafka Streams in Action teaches you everything you need to know to implement stream processing on data flowing into your Kafka platform, allowing you to focus on getting more from your data without sacrificing time or effort.

Big Data

Big Data teaches you to build big data systems using an architecture that takes advantage of clustered hardware along with new tools designed specifically to capture and analyze web-scale data.

Spark in Action

& Spark in Action 2nd Ed. - Spark in Action teaches you the theory and skills you need to effectively handle batch and streaming data using Spark. Fully updated for Spark 2.0.

In 3 lists

Kafka in Action

Kafka in Action is a fast-paced introduction to every aspect of working with Kafka you need to really reap its benefits.

Fusion in Action

Fusion in Action teaches you to build a full-featured data analytics pipeline, including document and data search and distributed data clustering.

Reactive Data Handling

Reactive Data Handling is a collection of five hand-picked chapters, selected by Manuel Bernhardt, that introduce you to building reactive applications capable of handling real-time processing with large data loads--free eBook!

Azure Data Engineering

A book about data engineering in general and the Azure platform specifically

Grokking Streaming Systems

Grokking Streaming Systems helps you unravel what streaming systems are, how they work, and whether they’re right for your business. Written to be tool-agnostic, you’ll be able to apply what you learn no matter which framework you choose.

Data Analysis with Python and PySpark

tutorial for using PySpark to build data-driven applications at scale.

In 2 lists

Data Pipelines with Apache Airflow

practical guide to building and maintaining data pipelines with Airflow.

In 2 lists

Distributed Systems for fun and profit

Theory of distributed systems. Include parts about time and ordering, replication and impossibility results.

In 3 lists

Graph-Powered Machine Learning

Alessandro Negro. Combine graph theory and models to improve machine learning projects

Books >Data Visualization

The beauty of data visualization

Designing Data Visualizations with Noah Iliinsky

Hans Rosling's 200 Countries, 200 Years, 4 Minutes

Ice Bucket Challenge Data Visualization

See category
87

Awesome Data Engineering

igorbarinov/awesome-data-engineering

A curated list of data engineering tools for software developers

Fresh★ 9.1k313 entriesPushed 22 days ago
84

Awesome Public Datasets

awesomedata/awesome-public-datasets

A topic-centric list of HQ open datasets.

Fresh★ 79k3 entriesPushed today
83

Awesome Ada

ohenley/awesome-ada

A curated list of awesome resources related to the Ada and SPARK programming language

Fresh★ 869418 entriesPushed 7 days ago
80

Awesome Network Analysis

briatte/awesome-network-analysis

A curated list of awesome network analysis resources.

Fresh★ 4.1k717 entriesPushed 1 month ago
78

Awesome Public Real-Time Datasets and Sources

bytewax/awesome-public-real-time-datasets

A list of publicly available datasets with real-time data maintained by the team at bytewax.io

Active★ 2.9k91 entriesPushed 2 months ago
74

Awesome Streaming

manuzhang/awesome-streaming

a curated list of awesome streaming frameworks, applications, etc

Fresh★ 3k1 entriesPushed today