Skip to content

Entry

LM Evaluation Harness

Appears in 7 awesome lists

Language Model Evaluation Harness is a framework to test generative language models on a large number of different evaluation tasks.

Open github.comeleutherai/lm-evaluation-harness

Found in these lists

Awesome LLM Resources

Section: 评估 Evaluation · A framework for few-shot evaluation of language models.

FreshScore 87

awesome-nlp

Section: Evaluation and Benchmarks · unified framework for LM benchmark evaluation.

FreshScore 90

Awesome Open Source AI

Section: 9. Evaluation, Benchmarks & Datasets · De-facto standard for generative model evaluation.

FreshScore 89

Awesome Production Machine Learning

Section: Evaluation and Monitoring · Language Model Evaluation Harness is a framework to test generative language models on a large number of different evaluation tasks.

FreshScore 92

Awesome Prompts

Section: Tools & Libraries · EleutherAI's unified LLM evaluation framework

FreshScore 90

awesome-python

Section: Other · A framework for few-shot evaluation of language models.

FreshScore 81

Indie Hacker Tools Plus

Section: LLM 评测与治理 (LLM Evaluation & Harness)

FreshScore 86

DeepEval

The most complete open-source LLM/agent eval framework: 20+ built-in metrics (hallucination, answer relevancy, RAGAs, tool correctness), pytest integration, and a CI-friendly runner. Removes the need to hand-roll eval infrastructure when you need structured, repeatable agent quality gates.

In 9 listsDetails

OpenAI Evals

Framework for evaluating LLMs and LLM systems with an open-source registry of 100+ community-contributed benchmarks. MIT licensed.

In 8 listsDetails

HuggingFace Datasets

The largest hub of ready-to-use NLP datasets for ML models with fast, easy-to-use and efficient data manipulation tools.

In 6 listsDetails

AutoRAG

RAG AutoML tool for automatically finding optimal RAG pipelines. Evaluates and optimizes retrieval-augmented generation with AutoML-style automation for your own data and use-case. Apache 2.0 licensed.

In 5 listsDetails

OpenCompass

OpenCompass is an LLM evaluation platform, supporting a wide range of models (Llama3, Mistral, InternLM2,GPT-4,LLaMa2, Qwen,GLM, Claude, etc) over 100+ datasets.

In 4 listsDetails

VLMEvalKit

Open-source evaluation toolkit for large multi-modality models (LMMs). Supports 220+ LMMs and 80+ benchmarks including MMMU, MathVista, and ChartQA. Powers the OpenVLM Leaderboard. Apache 2.0 licensed.

In 4 listsDetails

Evaluate

Evaluate machine learning models (huggingface). pycm - Multi-class confusion matrix. pandas_ml - Confusion matrix. Plotting learning curve: link. yellowbrick - Learning curve. pyroc - Receiver Operating Characteristic (ROC) curves.

In 4 listsDetails

OrcaPromptVault

Open corpus of the instructions and tool schemas shipping AI agents are actually sent: 119 artifacts from 43 products, 44 of them recorded off the wire with the command that reproduces each, every file marked captured or reported. AGPL-3.0.

In 4 listsDetails