Skip to content

Entry

Pathway

Appears in 8 awesome lists

Python ETL framework for stream processing, real-time analytics, LLM pipelines, and RAG. Features 350+ connectors with always-in-sync data from SharePoint, Google Drive, S3, Kafka, PostgreSQL and more. BSL 1.1 license (becomes Apache 2.0 after 4 years).

Open github.compathwaycom/pathway

Found in these lists

Awesome Ai Agents 2026

Section: RAG and Knowledge Bases · Live data RAG. Real-time streaming. 50k+ stars.

ActiveScore 74

Awesome Data Engineering

Section: Stream Processing · Performant open-source Python ETL framework with Rust runtime, supporting 300+ data sources.

FreshScore 87

Awesome local LLM

Section: Retrieval-Augmented Generation · Python ETL framework for stream processing, real-time analytics, LLM pipelines and RAG

FreshScore 87

Awesome Open Source AI

Section: 5. Retrieval-Augmented Generation (RAG) & Knowledge · Python ETL framework for stream processing, real-time analytics, LLM pipelines, and RAG. Features 350+ connectors with always-in-sync data from SharePoint, Google Drive, S3, Kafka, PostgreSQL and more. BSL 1.1 license (becomes Apache 2.0 after 4 years).

FreshScore 89

Awesome Pipeline

Section: Extract, transform, load (ETL) · Performant open-source Python ETL framework with Rust runtime, supporting 300+ data sources.

FreshScore 86

awesome-python

Section: Data Science and Analytics · Python ETL framework for stream processing, real-time analytics, LLM pipelines, and RAG.

FreshScore 81

Awesome Rust

Section: Data processing · Performant open-source Python ETL framework with Rust runtime, supporting 300+ data sources.

FreshScore 94

Awesome Python

Section: Data Ingestion / ETL · Python ETL framework for stream processing, real-time analytics, LLM pipelines, and RAG.

FreshScore 94

Milvus

Milvus is a cloud-native, open-source vector database built to manage embedding vectors generated by machine learning models and neural networks.

In 14 listsDetails

Mem0

Mem0 is an intelligent memory layer for Large Language Models that enhances personalized AI experiences by retaining and utilizing contextual information across various applications. github | website | docs | discord | twitter | github profile | linkedin

In 13 listsDetails

Haystack

Open-source AI orchestration framework for building context-engineered, production-ready LLM applications. Design modular pipelines and agent workflows with explicit control over retrieval, routing, memory, and generation. Built for scalable agents, RAG, multimodal applications, semantic search,…

In 13 listsDetails

Qdrant

Vector Search Engine and Database for the next generation of AI applications. Also available in the cloud

In 8 listsDetails

Chroma

An open-source embedding database for building AI applications with embeddings and semantic search.

In 8 listsDetails

Pinecone

The Pinecone vector database makes it easy to build high-performance vector search applications. Developer-friendly, fully managed, and easily scalable without infrastructure hassles.

In 8 listsDetails

Weaviate

An open source vector database that stores both objects and vectors, allowing for combining vector search with structured filtering with the fault-tolerance and scalability of a cloud-native database, all accessible through GraphQL, REST, and various language clients.

In 8 listsDetails

RAGFlow

RAGFlow is a leading open-source Retrieval-Augmented Generation (RAG) engine that fuses cutting-edge RAG with Agent capabilities to create a superior context layer for LLMs

In 7 listsDetails