Skip to content

Entry

Evaluate

Appears in 4 awesome lists

Evaluate machine learning models (huggingface). pycm - Multi-class confusion matrix. pandas_ml - Confusion matrix. Plotting learning curve: link. yellowbrick - Learning curve. pyroc - Receiver Operating Characteristic (ROC) curves.

Open github.comhuggingface/evaluate

Found in these lists

awesome-nlp

Section: Libraries · reference implementations for NLP metrics.

FreshScore 90

Awesome Open Source AI

Section: 9. Evaluation, Benchmarks & Datasets · Standardized evaluation metrics.

FreshScore 89

Awesome Production Machine Learning

Section: Evaluation and Monitoring · Evaluate is a library that makes evaluating and comparing models and reporting their performance easier and more standardized.

FreshScore 92

Awesome Data Science with Python

Section: General · Evaluate machine learning models (huggingface). pycm - Multi-class confusion matrix. pandas_ml - Confusion matrix. Plotting learning curve: link. yellowbrick - Learning curve. pyroc - Receiver Operating Characteristic (ROC) curves.

FreshScore 82

Opik

Comet's open-source AI observability and evaluation platform: deep tracing of LLM calls, conversation logging, and agent activity, plus built-in eval metrics, prompt versioning, guardrails, and the Opik Agent Optimizer. Worth including because it unifies observability, verification, and…

In 16 listsDetails

transformers

(formerly known as pytorch-transformers and pytorch-pretrained-bert) provides state-of-the-art general-purpose architectures (BERT, GPT-2, RoBERTa, XLM, DistilBert, XLNet, CTRL...) for Natural Language Understanding (NLU) and Natural Language Generation (NLG) with over 32+ pretrained models in…

In 14 listsDetails

Haystack

Open-source AI orchestration framework for building context-engineered, production-ready LLM applications. Design modular pipelines and agent workflows with explicit control over retrieval, routing, memory, and generation. Built for scalable agents, RAG, multimodal applications, semantic search,…

In 13 listsDetails

promptfoo

Test your prompts, models, RAGs. Evaluate and compare LLM outputs, catch regressions, and improve prompt quality. LLM evals for OpenAI/Azure GPT, Anthropic Claude, VertexAI Gemini, Ollama, Local & private models like Mistral/Mixtral/Llama with CI/CD

In 11 listsDetails

Awesome NLP with Ruby

A curated list of resources dedicated to Natural Language Processing and text processing for Ruby.

In 5 listsDetails

Phoenix

Open-source AI observability & evaluation platform (Arize) — OpenTelemetry-native tracing for agents, LLM-as-judge evals, versioned datasets & experiments for prompt regression testing, prompt management with version control and replay, plus an MCP endpoint so Claude Code/Cursor can query traces…

In 10 listsDetails

Langfuse

The most widely adopted self-hostable LLM observability platform: traces every agent step, manages prompt versions, and runs evals in one tool. Preferred over cloud-only alternatives when data residency or cost control is a constraint.

In 10 listsDetails

DeepEval

The most complete open-source LLM/agent eval framework: 20+ built-in metrics (hallucination, answer relevancy, RAGAs, tool correctness), pytest integration, and a CI-friendly runner. Removes the need to hand-roll eval infrastructure when you need structured, repeatable agent quality gates.

In 9 listsDetails