Skip to content

Entry

Triton Inference Server

Appears in 5 awesome lists

NVIDIA's production-grade open-source inference serving software. Supports multiple frameworks (TensorRT, PyTorch, ONNX) with optimized cloud and edge deployment.

Open github.comtriton-inference-server/server

Found in these lists

Awesome LLMOps

Section: Frameworks/Servers for Serving · The Triton Inference Server provides an optimized cloud and edge inferencing solution.

ActiveScore 75

Awesome MLOps

Section: Model Serving · Provides an optimized cloud and edge inferencing solution.

FreshScore 80

Awesome Open Source AI

Section: 3. Inference Engines & Serving · NVIDIA's production-grade open-source inference serving software. Supports multiple frameworks (TensorRT, PyTorch, ONNX) with optimized cloud and edge deployment.

FreshScore 89

Awesome Production Machine Learning

Section: Deployment and Serving · Triton is a high performance open source serving software to deploy AI models from any framework on GPU & CPU while maximizing utilization.

FreshScore 92

awesome-python

Section: Other · The Triton Inference Server provides an optimized cloud and edge inferencing solution.

FreshScore 81

LocalAI

robot: The free, Open Source alternative to OpenAI, Claude and others. Self-hosted and local-first. Drop-in replacement for OpenAI, running on consumer-grade hardware. No GPU required. Runs gguf, transformers, diffusers and many more models architectures. Features: Generate Text, Audio, Video,…

In 14 listsDetails

Ollama

Ollama is a tool for running large language models locally, offering easy setup for macOS, Windows, Linux, and Docker, along with a library of models and quickstart guides for customization and integration github | github profile

In 12 listsDetails

KubeStellar Console

Multi-cluster Kubernetes MCP server bridging Gemini CLI to kubeconfig and Kubernetes APIs. Manage clusters, policies, and 20+ CNCF project integrations across edge and cloud. Install via brew tap kubestellar/tap && brew install kc-agent.

In 13 listsDetails

Gradio

Build and share delightful machine learning apps, all in Python. The de facto standard for creating interactive ML demos with automatic UI generation from function signatures. Powers thousands of Hugging Face Spaces.

In 11 listsDetails

vLLM

State-of-the-art serving engine with PagedAttention and continuous batching. Currently the fastest production-grade LLM server.

In 11 listsDetails

Streamlit

The fastest way to build and share data apps. Transform Python scripts into beautiful web applications with minimal code. Widely used for ML model demos, data visualization, and internal tools.

In 10 listsDetails

llama.cpp

Pure C/C++ inference engine with GGUF format support. The gold standard for CPU/GPU/Apple Silicon on-device running. Includes llama-server for OpenAI-compatible API. Now at 100K+ stars.

In 10 listsDetails

OpenLLM

Production-grade platform for running any open-source LLMs as OpenAI-compatible API endpoints. Supports 50+ models with built-in streaming, batching, and auto-acceleration. Apache 2.0 licensed.

In 10 listsDetails