Skip to content
83

Awesome Model Quantization

A curated collection of papers, benchmarks, surveys, and tools for model quantization, covering low-bit networks, LLMs, multimodal and generative models, vector and lattice quantization, and efficient deployment.

2.5k stars241 forks35 entriesLast push Sep 28, 2026 (2 days ago)License none

This page lists names, links and short descriptions. The original list on GitHub is the source and belongs to its authors.

Benchmarks

MQBench: Towards Reproducible and Deployable Model Quantization Benchmark

QAT + deploymentCompares quantization algorithms under reproducible settings and hardware backend constraints.

BiBench: Benchmarking and Analyzing Network Binarization

Binary networksCompares binarization methods across tasks, architectures and deployment settings.

Evaluating Quantized Large Language Models

Weights, activations + KV cacheEvaluates 11 model families on basic NLP, emergent abilities, trustworthiness, dialogue and long-context tasks.

LLMC: Benchmarking Large Language Model Quantization with a Versatile Compression Toolkit

LLM toolkitCompares calibration data, method pipelines and quantization configurations; the toolkit is now LightCompress.

An empirical study of LLaMA3 quantization: from LLMs to MLLMs

LLMs + multimodalExamines low-bit behavior across LLaMA3 language and multimodal models.

An Empirical Study of Qwen3 Quantization

Dense + MoE LLMsStudies quantization across Qwen3 model sizes, architectures and reasoning settings.

RobustMQ: Benchmarking Robustness of Quantized Models

Model robustnessTests quantized models beyond clean accuracy, including robustness under input perturbations.

Qwen3 study

Preprint

Survey Papers

A White Paper on Neural Network Quantization

Practical PTQ + QATExplains quantizer design, common failure modes and practical post-training and quantization-aware training workflows.

A Survey of Quantization Methods for Efficient Neural Network Inference

Foundations + taxonomyReviews quantization design choices, mixed precision and the trade-offs between model accuracy and efficient inference.

Binary Neural Networks: A Survey

Binary networksSurveys binary network representations, training methods and applications.

A Survey of Low-bit Large Language Models: Basics, Systems, and Algorithms

LLM algorithms + systemsConnects low-bit LLM algorithms with numerical formats and inference systems.

Low-bit Model Quantization for Deep Neural Networks: A Survey

Broad low-bit methodsMaps low-bit quantization methods across neural network architectures and applications.

Quantization methods survey

Chapter

Binary networks survey

Blog

Books

Quantization and Fast Inference: A practitioner’s guide to efficient AI

Vivek KalyanaranganManning, early access (MEAP)

Efficient Processing of Deep Neural Networks

Vivienne Sze, Yu-Hsin Chen, Tien-Ju Yang, Joel S. Emer2020

Vector Quantization and Signal Compression

Allen Gersho, Robert M. Gray1992

Machine Learning Systems

Vijay Janapa Reddi and contributorsOpen-access online textbook

In 2 lists

Related Repositories >Quantization and training toolkits

TorchAO

PyTorch-native quantization for training and inference.

In 2 lists

bitsandbytes

Low-bit linear layers and quantized optimizers, including implementations used by LLM.int8() and QLoRA.

In 4 listsDetails

LLM Compressor

Model compression and quantization workflows for deployment with vLLM.

In 2 lists

NVIDIA Model Optimizer

Quantization and model optimization with export to supported inference runtimes.

In 2 lists

LightCompress (formerly LLMC)

Research and deployment toolkit spanning LLMs, vision-language and generative models.

HQQ

Half-quadratic weight quantization without calibration data.

AIMET

Post-training and quantization-aware model optimization.

Brevitas

PyTorch quantization-aware training with configurable quantizers and hardware export.

Related Repositories >Inference and hardware

llama.cpp

Local LLM inference with GGUF models and multiple quantization formats.

In 10 listsDetails

vLLM

LLM serving with supported low-bit kernels and quantized KV caches.

In 11 listsDetails

TensorRT LLM

NVIDIA GPU inference with supported low-precision formats and optimized kernels.

In 7 listsDetails

Transformer Engine

Low-precision transformer computation, including FP8 and FP4 on supported NVIDIA GPUs.

Nunchaku

Low-bit diffusion inference, including SVDQuant kernels.

BitNet

Inference framework for supported native low-bit BitNet models.

In 5 listsDetails

FINN

Dataflow compilation for quantized neural networks on FPGAs.

In 2 lists

ncnn

Mobile neural network inference, including INT8 deployment.

In 6 listsDetails
See category
94

Awesome OpenClaw Skills

VoltAgent/awesome-openclaw-skills

The awesome collection of OpenClaw skills. 5,400+ skills filtered and categorized from the official OpenClaw Skills Registry.🦞

Fresh★ 53k830 entriesPushed today
92

Awesome DeepSeek Harness (DSH) Plugin

awesome-dsh-plugin/awesome-dsh-plugin

A curated list of plugins for DeepSeek Harness (dsh) · DeepSeek Harness 插件精选列表

Fresh★ 17k1654 entriesPushed today
91

Awesome Guidelines

Kristories/awesome-guidelines

Programming style, best practices, and coding conventions.

Fresh★ 11k166 entriesPushed 2 days ago
90

Awesome

sindresorhus/awesome

😎 Awesome lists about all kinds of interesting topics [NOTE: Pull requests are temporarily disabled until I have a chance to catch up with the existing ones]

Fresh★ 513k51 entriesPushed 28 days ago
90

Awesome Prompts

ai-boost/awesome-prompts

Curated list of chatgpt prompts from the top-rated GPTs in the GPTs Store. Prompt Engineering, prompt attack & prompt protect. Advanced Prompt Engineering papers.

Fresh★ 9k288 entriesPushed yesterday
90

Awesome README

matiassingers/awesome-readme

A curated list of awesome READMEs

Fresh★ 22k143 entriesPushed yesterday