Awesome Artificial Intelligence
Section: Evals and reliability · An open-source framework and registry for evaluating language models and systems.
Entry
Appears in 8 awesome lists
Framework for evaluating LLMs and LLM systems with an open-source registry of 100+ community-contributed benchmarks. MIT licensed.
Section: Evals and reliability · An open-source framework and registry for evaluating language models and systems.
Section: Openai · Evals is a framework for evaluating OpenAI models and an open-source registry of benchmarks.
Section: LLM Evaluation · Evals is a framework for evaluating LLMs and LLM systems, and an open-source registry of benchmarks.
Section: 9. Evaluation, Benchmarks & Datasets · Framework for evaluating LLMs and LLM systems with an open-source registry of 100+ community-contributed benchmarks. MIT licensed.
Section: Evaluation and Monitoring · Evals is a framework for evaluating OpenAI models and an open-source registry of benchmarks.
Section: Eval & Testing · Open eval framework and benchmark registry — standardizes LLM performance measurement.
Section: Other · Evals is a framework for evaluating LLMs and LLM systems, and an open-source registry of benchmarks.
Section: Evaluation and Observability · Framework and registry for benchmarking model behavior.
industrial-grade ASR toolkit; 170× realtime on GPU, 50+ languages, built-in VAD, punctuation, speaker diarization, and emotion detection. Includes non-autoregressive SenseVoice and LLM-based Fun-ASR-Nano models.
AutoGPT provides accessible AI tools for building and using AI agents, offering a comprehensive framework including Forge for agent creation, agbenchmark for performance evaluation, a leaderboard for competition, a user-friendly UI, and CLI for seamless integration and management github | github…
The most complete open-source LLM/agent eval framework: 20+ built-in metrics (hallucination, answer relevancy, RAGAs, tool correctness), pytest integration, and a CI-friendly runner. Removes the need to hand-roll eval infrastructure when you need structured, repeatable agent quality gates.
OpenAI's cookbook includes examples of prompt engineering.
onWatch is a lightweight Go CLI that tracks AI API quota usage across multiple providers (Anthropic Pro/Max Plans, Codex, Gemini CLI, Synthetic, Z.ai, GitHub Copilot, MiniMax Coding/Token Plan, Antigravity, OpenRouter) in real time, with consumption rate projections, historical usage graphs, and…