Skip to content

Entry

OpenAI Evals

Appears in 8 awesome lists

Framework for evaluating LLMs and LLM systems with an open-source registry of 100+ community-contributed benchmarks. MIT licensed.

Open github.comopenai/evals

Found in these lists

Awesome Artificial Intelligence

Section: Evals and reliability · An open-source framework and registry for evaluating language models and systems.

FreshScore 86

awesome-ChatGPT-repositories

Section: Openai · Evals is a framework for evaluating OpenAI models and an open-source registry of benchmarks.

FreshScore 87

Awesome Generative AI

Section: LLM Evaluation · Evals is a framework for evaluating LLMs and LLM systems, and an open-source registry of benchmarks.

SlowScore 62

Awesome Open Source AI

Section: 9. Evaluation, Benchmarks & Datasets · Framework for evaluating LLMs and LLM systems with an open-source registry of 100+ community-contributed benchmarks. MIT licensed.

FreshScore 89

Awesome Production Machine Learning

Section: Evaluation and Monitoring · Evals is a framework for evaluating OpenAI models and an open-source registry of benchmarks.

FreshScore 92

Awesome Prompts

Section: Eval & Testing · Open eval framework and benchmark registry — standardizes LLM performance measurement.

FreshScore 90

awesome-python

Section: Other · Evals is a framework for evaluating LLMs and LLM systems, and an open-source registry of benchmarks.

FreshScore 81

Awesome Startup

Section: Evaluation and Observability · Framework and registry for benchmarking model behavior.

FreshScore 82

FunASR

industrial-grade ASR toolkit; 170× realtime on GPU, 50+ languages, built-in VAD, punctuation, speaker diarization, and emotion detection. Includes non-autoregressive SenseVoice and LLM-based Fun-ASR-Nano models.

In 9 listsDetails

AutoGPT

AutoGPT provides accessible AI tools for building and using AI agents, offering a comprehensive framework including Forge for agent creation, agbenchmark for performance evaluation, a leaderboard for competition, a user-friendly UI, and CLI for seamless integration and management github | github…

In 10 listsDetails

DeepEval

The most complete open-source LLM/agent eval framework: 20+ built-in metrics (hallucination, answer relevancy, RAGAs, tool correctness), pytest integration, and a CI-friendly runner. Removes the need to hand-roll eval infrastructure when you need structured, repeatable agent quality gates.

In 9 listsDetails

OpenAI Cookbook

OpenAI's cookbook includes examples of prompt engineering.

In 9 listsDetails

go-openai

OpenAI ChatGPT, GPT-3, GPT-4, DALL·E, Whisper API wrapper for Go

In 6 listsDetails

onWatch

onWatch is a lightweight Go CLI that tracks AI API quota usage across multiple providers (Anthropic Pro/Max Plans, Codex, Gemini CLI, Synthetic, Z.ai, GitHub Copilot, MiniMax Coding/Token Plan, Antigravity, OpenRouter) in real time, with consumption rate projections, historical usage graphs, and…

In 6 listsDetails

Auto-GPT

An experimental open-source attempt to make GPT-4 fully autonomous.

In 7 listsDetails

ToolJet

Low-code platform for building business applications. Connect to databases, cloud storages, GraphQL, API endpoints, Airtable, Google sheets, OpenAI, etc and build apps using drag and drop application builder. Built using JavaScript/TypeScript. 🚀

In 6 listsDetails