Skip to content

Entry

Trino

Appears in 4 awesome lists

Distributed SQL query engine designed to query large data sets distributed over one or more heterogeneous data sources.

Open github.comtrinodb/trino

Found in these lists

Awesome Data Analysis

Section: Tools · A distributed SQL query engine designed for fast analytic queries against large datasets.

FreshScore 80

Awesome Database Tools

Section: Über SQL · Distributed SQL query engine designed to query large data sets distributed over one or more heterogeneous data sources.

ActiveScore 75

Awesome First Pull Request Opportunities

Section: Java · (label: good first issue) A distributed SQL query engine for big data. Ask for guidance on project's Slack.

ActiveScore 83

Awesome Java

Section: Database · Apache-2.0 🟢Distributed SQL query engine for big data.

FreshScore 93

Luigi

Python module for building complex pipelines of batch jobs. Handles dependency resolution, workflow management, visualization, and Hadoop integration. Built at Spotify and battle-tested in production. Apache 2.0 licensed.

In 15 listsDetails

Apache Airflow

"Use airflow to author workflows as directed acyclic graphs (DAGs) of tasks. The airflow scheduler executes your tasks on an array of workers while following the specified dependencies. Rich command line utilities make performing complex surgeries on DAGs a snap. The rich user interface makes it…

In 13 listsDetails

Dagster

Cloud-native orchestration platform for developing and maintaining data assets including ML models. Declarative programming model with integrated lineage and observability. Apache 2.0 licensed.

In 12 listsDetails

Prefect

Workflow management system that makes it easy to take your data pipelines and add semantics like retries, logging, dynamic mapping, caching, failure notifications, and more.

In 11 listsDetails

Apache Spark

Unified analytics engine for large-scale data processing. In-memory cluster computing with high-level APIs in Python, Scala, Java, and R. Powers MLlib for distributed machine learning and Structured Streaming for real-time data. Apache 2.0 licensed.

In 8 listsDetails

Kestra

Event-driven orchestration and scheduling platform for mission-critical workflows. Infrastructure-as-Code approach with declarative YAML, Git version control integration, and hundreds of plugins for data pipelines and ML workflows. Apache 2.0 licensed.

In 8 listsDetails

Elasticsearch

Distributed search and analytics engine with native k-NN vector search, hybrid search, and dense vector indexing. Industry-standard for full-text search now with powerful semantic search capabilities. AGPL-3.0/Elastic-2.0 dual licensed.

In 7 listsDetails

Apache Flink

Stream processing framework with powerful batch and streaming capabilities. High-throughput, low-latency runtime with exactly-once processing guarantees. Ideal for real-time AI inference pipelines and event-driven ML applications. Apache 2.0 licensed.

In 7 listsDetails