Skip to content

Entry

Apache Hudi

Appears in 4 awesome lists

Hudi is a transactional data lake platform that brings core warehouse and database functionality directly to a data lake. Hudi is great for streaming workloads, and also allows creation of efficient incremental batch pipelines. Supports popular query engines including Spark, Flink, Presto, Trino,…

Open github.comapache/hudi

Found in these lists

Awesome Data Analysis

Section: Tools · An open data lakehouse platform, built on a high-performance open table format.

FreshScore 80

Awesome Open Source AI

Section: 1. Core Frameworks & Libraries · Open data lakehouse platform for ingesting, indexing, storing, serving, transforming and managing data across cloud environments. Supports upserts, deletes and incremental processing on big data with built-in ingestion tools for Spark and Flink. Apache 2.0 licensed.

FreshScore 89

Awesome Production Machine Learning

Section: Data Storage Optimisation · Hudi is a transactional data lake platform that brings core warehouse and database functionality directly to a data lake. Hudi is great for streaming workloads, and also allows creation of efficient incremental batch pipelines. Supports popular query engines including Spark, Flink, Presto, Trino,…

FreshScore 92

Awesome Spark

Section: Storage · Upserts, Deletes And Incremental Processing on Big Data..

SlowScore 66

TensorFlow

How to use the Hexagon Delegate to speed up model inference on mobile and edge devices. Also see blog post Accelerating TensorFlow Lite on Qualcomm Hexagon DSPs.

In 23 listsDetails

PyTorch

(label: good first issue) PyTorch is an open source machine learning library based on the Torch library, used for applications such as computer vision and natural language processing.

In 16 listsDetails

Luigi

Python module for building complex pipelines of batch jobs. Handles dependency resolution, workflow management, visualization, and Hadoop integration. Built at Spotify and battle-tested in production. Apache 2.0 licensed.

In 15 listsDetails

transformers

(formerly known as pytorch-transformers and pytorch-pretrained-bert) provides state-of-the-art general-purpose architectures (BERT, GPT-2, RoBERTa, XLM, DistilBert, XLNet, CTRL...) for Natural Language Understanding (NLU) and Natural Language Generation (NLG) with over 32+ pretrained models in…

In 14 listsDetails

Milvus

Milvus is a cloud-native, open-source vector database built to manage embedding vectors generated by machine learning models and neural networks.

In 14 listsDetails

Apache Airflow

"Use airflow to author workflows as directed acyclic graphs (DAGs) of tasks. The airflow scheduler executes your tasks on an array of workers while following the specified dependencies. Rich command line utilities make performing complex surgeries on DAGs a snap. The rich user interface makes it…

In 13 listsDetails

Dagster

Cloud-native orchestration platform for developing and maintaining data assets including ML models. Declarative programming model with integrated lineage and observability. Apache 2.0 licensed.

In 12 listsDetails

XGBoost

Scalable, Portable and Distributed Gradient Boosting (GBDT, GBRT or GBM) Library, for Python, R, Java, Scala, C++ and more. Runs on single machine, Hadoop, Spark, Flink and DataFlow. [Apache2]

In 11 listsDetails