Awesome Ai Agents 2026
Section: Tracing and Monitoring · OSS LLM observability. Traces, evals, prompts.
Entry
Appears in 10 awesome lists
The most widely adopted self-hostable LLM observability platform: traces every agent step, manages prompt versions, and runs evals in one tool. Preferred over cloud-only alternatives when data residency or cost control is a constraint.
Section: Tracing and Monitoring · OSS LLM observability. Traces, evals, prompts.
Section: Links
Section: Prompts · 🪢 Open source LLM engineering platform: LLM Observability, metrics, evals, prompt management, playground, datasets. Integrates with LlamaIndex, Langchain, OpenAI SDK, LiteLLM, and more. 🍊YC W23
Section: Observability & Tracing · The most widely adopted self-hostable LLM observability platform: traces every agent step, manages prompt versions, and runs evals in one tool. Preferred over cloud-only alternatives when data residency or cost control is a constraint.
Section: LLMOps · Open Source LLM Engineering Platform: Traces, evals, prompt management and metrics to debug and improve your LLM application.
Section: Testing, Evaluation and Observability · an open-source LLM engineering platform: LLM Observability, metrics, evals, prompt management, playground, datasets. Integrates with OpenTelemetry, Langchain, OpenAI SDK, LiteLLM, and more
Section: 8. MLOps / LLMOps & Production · #1 open-source LLM observability platform.
Section: Evaluation and Monitoring · Langfuse is an observability & analytics solution for LLM-based applications.
Section: Eval & Observability · Open-source LLM engineering platform — tracing, evals, prompt management, A/B experiments.
Section: Repositories · Langfuse, an open-source LLM engineering platform, offers debugging, prompt management, metrics for LLM apps improvement, and won the #1 Golden Kitty in the AI Infra Category from Product Hunt github | website | twitter | discord
Test your prompts, models, RAGs. Evaluate and compare LLM outputs, catch regressions, and improve prompt quality. LLM evals for OpenAI/Azure GPT, Anthropic Claude, VertexAI Gemini, Ollama, Local & private models like Mistral/Mixtral/Llama with CI/CD
LLM engineering platform for model tracing, prompt management, and application evaluation. Langfuse helps teams collaboratively debug, analyze, and iterate on their LLM applications such as chatbots or AI agents. (Demo, Source Code, Clients) MIT Docker
🐙 Guides, papers, lecture, notebooks and resources for prompt engineering
Open-source AI observability & evaluation platform (Arize) — OpenTelemetry-native tracing for agents, LLM-as-judge evals, versioned datasets & experiments for prompt regression testing, prompt management with version control and replay, plus an MCP endpoint so Claude Code/Cursor can query traces…
Private AI platform for building intelligent agents and assistants with enterprise search. Features Agent Builder, deep research tools, multi-format document analysis, and multi-model support. MIT licensed.
DemoGPT enables you to create quick demos by just using prompt. It applies ToT approach on Langchain documentation tree.
Open-source LLM observability proxy (YC W23) with the largest open-source pricing database (300+ models). One-line proxy integration provides cost tracking, token monitoring, session tracing, and prompt versioning across providers. The AI Gateway component handles request routing and caching with…
🟢🟠 — Open AI gateway with provider routing, fallback and retry controls, guardrail integrations, observability, and MCP traffic support for model and agent applications. (Portkey) — note: the gateway is general infrastructure rather than a standalone security scanner; model-provider…