Awesome Data Engineering
Section: Stream Processing · An open source ETL framework to build fresh index for AI.
Entry
Appears in 6 awesome lists
ETL framework to build fresh context for AI agents, with incremental processing
Section: Stream Processing · An open source ETL framework to build fresh index for AI.
Section: Pipeline frameworks & libraries · ETL framework to build fresh index.
Section: LLM and Inference · Incremental engine for long horizon agents Star if you like it!
Section: Library
Section: CocoIndex (24 · 6.1K) - Data transformation framework for AI. Ultra performant, with.. Apache-2 · (👨💻 93 · 🔀 450 · 📦 73):
Section: MLOps · ETL framework to build fresh context for AI agents, with incremental processing
Langchain integrates various providers like Anthropic, AWS, and OpenAI, and offers tools for components such as LLMs, chat models, and data analysis, supporting functionalities from Alpha Vantage to YouTube github | docs
Unified proxy and SDK that routes to 100+ LLM providers behind a single OpenAI-compatible interface, with a Router handling retry/fallback across deployments, per-project cost and rate-limit tracking, and OTEL callback integrations. The right infrastructure layer when your harness needs provider…
Python module for building complex pipelines of batch jobs. Handles dependency resolution, workflow management, visualization, and Hadoop integration. Built at Spotify and battle-tested in production. Apache 2.0 licensed.
February 2026 release making human oversight a native workflow primitive: suspend execution at critical decision points, expose review-and-edit UI mid-flow, and route subsequent execution based on human action (approve/reject/escalate). Demonstrates how HITL transitions from bolt-on approval gates…
Mem0 is an intelligent memory layer for Large Language Models that enhances personalized AI experiences by retaining and utilizing contextual information across various applications. github | website | docs | discord | twitter | github profile | linkedin
A fast and simple framework for building and running distributed applications. Ray is packaged with RLlib, a scalable reinforcement learning library, and Tune, a scalable hyperparameter tuning library. ray.io
Open-source AI orchestration framework for building context-engineered, production-ready LLM applications. Design modular pipelines and agent workflows with explicit control over retrieval, routing, memory, and generation. Built for scalable agents, RAG, multimodal applications, semantic search,…
June 2026 harness-first redesign built around the Capability primitive: a single composable unit bundling instructions, tools, lifecycle hooks, and model settings. The split between a small stable core and a fast-moving pydantic-ai-harness lets capabilities graduate as they prove essential, while…