Skip to content

Entry

DuckDB

Appears in 8 awesome lists

Efficiently run SQL queries on pandas DataFrame, duckplyr for R, Great Intro. ducklake - Duckdb extention for storing data in a datalake. fireducks - Speedier alternative to pandas with similar API. pandasvault - Large collection of pandas tricks. polars - Multi-threaded alternative to pandas.…

Open github.comduckdb/duckdb

Found in these lists

Awesome Data Analysis

Section: Databases · In‑process analytical database designed for fast OLAP queries.

FreshScore 80

AWESOME DATA SCIENCE

Section: Miscellaneous Tools · An in-process SQL OLAP database management system

FreshScore 92

Awesome Open Source AI

Section: 1. Core Frameworks & Libraries · High-performance analytical in-process SQL database system. Fast, reliable, portable, and easy to use with rich SQL dialect support. Perfect for data processing and analytics workloads. MIT licensed.

FreshScore 89

Libraries and packages

Section: Databases · In-process analytical SQL database that queries Parquet and Arrow files directly, a common backend for research datasets

FreshScore 86

Awesome Data Science with Python

Section: General · Efficiently run SQL queries on pandas DataFrame, duckplyr for R, Great Intro. ducklake - Duckdb extention for storing data in a datalake. fireducks - Speedier alternative to pandas with similar API. pandasvault - Large collection of pandas tricks. polars - Multi-threaded alternative to pandas.…

FreshScore 82

awesome-cpp

Section: Other · DuckDB is an analytical in-process SQL database management system

FreshScore 79

Awesome Python

Section: Database · An in-process SQL OLAP database management system; optimized for analytics and fast queries, similar to SQLite but for analytical workloads.

FreshScore 94

Awesome Systematic Trading

Section: Databases · | C++, Python | - ArcticDB is a high performance, serverless DataFrame database built for the Python Data Science ecosystem.

FreshScore 89

TensorFlow

How to use the Hexagon Delegate to speed up model inference on mobile and edge devices. Also see blog post Accelerating TensorFlow Lite on Qualcomm Hexagon DSPs.

In 23 listsDetails

PyTorch

(label: good first issue) PyTorch is an open source machine learning library based on the Torch library, used for applications such as computer vision and natural language processing.

In 16 listsDetails

Opik

Comet's open-source AI observability and evaluation platform: deep tracing of LLM calls, conversation logging, and agent activity, plus built-in eval metrics, prompt versioning, guardrails, and the Opik Agent Optimizer. Worth including because it unifies observability, verification, and…

In 16 listsDetails

Luigi

Python module for building complex pipelines of batch jobs. Handles dependency resolution, workflow management, visualization, and Hadoop integration. Built at Spotify and battle-tested in production. Apache 2.0 licensed.

In 15 listsDetails

transformers

(formerly known as pytorch-transformers and pytorch-pretrained-bert) provides state-of-the-art general-purpose architectures (BERT, GPT-2, RoBERTa, XLM, DistilBert, XLNet, CTRL...) for Natural Language Understanding (NLU) and Natural Language Generation (NLG) with over 32+ pretrained models in…

In 14 listsDetails

Apache Airflow

"Use airflow to author workflows as directed acyclic graphs (DAGs) of tasks. The airflow scheduler executes your tasks on an array of workers while following the specified dependencies. Rich command line utilities make performing complex surgeries on DAGs a snap. The rich user interface makes it…

In 13 listsDetails

Ray

A fast and simple framework for building and running distributed applications. Ray is packaged with RLlib, a scalable reinforcement learning library, and Tune, a scalable hyperparameter tuning library. ray.io

In 13 listsDetails

XGBoost

Scalable, Portable and Distributed Gradient Boosting (GBDT, GBRT or GBM) Library, for Python, R, Java, Scala, C++ and more. Runs on single machine, Hadoop, Spark, Flink and DataFlow. [Apache2]

In 11 listsDetails