Skip to content

Entry

Apache Spark

Appears in 8 awesome lists

Unified analytics engine for large-scale data processing. In-memory cluster computing with high-level APIs in Python, Scala, Java, and R. Powers MLlib for distributed machine learning and Structured Streaming for real-time data. Apache 2.0 licensed.

Open github.comapache/spark

Found in these lists

Awesome Big Data

Section: SQL-like processing · is a Query Optimization Framework for Spark and Shark.

ActiveScore 84

Awesome Data Analysis

Section: Tools · A unified engine for large-scale data processing and analytics.

FreshScore 80

Awesome Integration

Section: Stream Processing · Unified analytics engine whose Structured Streaming API provides scalable, fault-tolerant stream processing on the Spark SQL engine.

FreshScore 82

Awesome Machine Learning

Section: Java · Spark is a fast and general engine for large-scale data processing.

FreshScore 93

Awesome Open Source AI

Section: 1. Core Frameworks & Libraries · Unified analytics engine for large-scale data processing. In-memory cluster computing with high-level APIs in Python, Scala, Java, and R. Powers MLlib for distributed machine learning and Structured Streaming for real-time data. Apache 2.0 licensed.

FreshScore 89

Awesome Production Machine Learning

Section: Data Stream Processing · Micro-batch processing for streams using the apache spark framework as a backend supporting stateful exactly-once semantics.

FreshScore 92

Awesome Python

Section: Distributed Computing · Apache Spark Python API.

FreshScore 94

Awesome Systematic Trading

Section: Computation · | Scala | - Apache Spark - A unified analytics engine for large-scale data processing

FreshScore 89

Luigi

Python module for building complex pipelines of batch jobs. Handles dependency resolution, workflow management, visualization, and Hadoop integration. Built at Spotify and battle-tested in production. Apache 2.0 licensed.

In 15 listsDetails

Apache Airflow

"Use airflow to author workflows as directed acyclic graphs (DAGs) of tasks. The airflow scheduler executes your tasks on an array of workers while following the specified dependencies. Rich command line utilities make performing complex surgeries on DAGs a snap. The rich user interface makes it…

In 13 listsDetails

Dagster

Cloud-native orchestration platform for developing and maintaining data assets including ML models. Declarative programming model with integrated lineage and observability. Apache 2.0 licensed.

In 12 listsDetails

Prefect

Workflow management system that makes it easy to take your data pipelines and add semantics like retries, logging, dynamic mapping, caching, failure notifications, and more.

In 11 listsDetails

DuckDB

A fast in-process analytical database that has zero external dependencies, runs on Linux/macOS/Windows, offers a rich SQL dialect, and is free and extensible.

In 8 listsDetails

Kestra

Event-driven orchestration and scheduling platform for mission-critical workflows. Infrastructure-as-Code approach with declarative YAML, Git version control integration, and hundreds of plugins for data pipelines and ML workflows. Apache 2.0 licensed.

In 8 listsDetails

Apache Flink

Stream processing framework with powerful batch and streaming capabilities. High-throughput, low-latency runtime with exactly-once processing guarantees. Ideal for real-time AI inference pipelines and event-driven ML applications. Apache 2.0 licensed.

In 7 listsDetails

Kedro

Toolbox for production-ready data science. Uses software engineering best practices to help you create data engineering and data science pipelines that are reproducible, maintainable, and modular. Apache 2.0 licensed.

In 7 listsDetails