Awesome LLMOps
Section: Frameworks/Servers for Serving · The Triton Inference Server provides an optimized cloud and edge inferencing solution.
Entry
Appears in 5 awesome lists
NVIDIA's production-grade open-source inference serving software. Supports multiple frameworks (TensorRT, PyTorch, ONNX) with optimized cloud and edge deployment.
Section: Frameworks/Servers for Serving · The Triton Inference Server provides an optimized cloud and edge inferencing solution.
Section: Model Serving · Provides an optimized cloud and edge inferencing solution.
Section: 3. Inference Engines & Serving · NVIDIA's production-grade open-source inference serving software. Supports multiple frameworks (TensorRT, PyTorch, ONNX) with optimized cloud and edge deployment.
Section: Deployment and Serving · Triton is a high performance open source serving software to deploy AI models from any framework on GPU & CPU while maximizing utilization.
Section: Other · The Triton Inference Server provides an optimized cloud and edge inferencing solution.
robot: The free, Open Source alternative to OpenAI, Claude and others. Self-hosted and local-first. Drop-in replacement for OpenAI, running on consumer-grade hardware. No GPU required. Runs gguf, transformers, diffusers and many more models architectures. Features: Generate Text, Audio, Video,…
Ollama is a tool for running large language models locally, offering easy setup for macOS, Windows, Linux, and Docker, along with a library of models and quickstart guides for customization and integration github | github profile
Multi-cluster Kubernetes MCP server bridging Gemini CLI to kubeconfig and Kubernetes APIs. Manage clusters, policies, and 20+ CNCF project integrations across edge and cloud. Install via brew tap kubestellar/tap && brew install kc-agent.
Build and share delightful machine learning apps, all in Python. The de facto standard for creating interactive ML demos with automatic UI generation from function signatures. Powers thousands of Hugging Face Spaces.
State-of-the-art serving engine with PagedAttention and continuous batching. Currently the fastest production-grade LLM server.
The fastest way to build and share data apps. Transform Python scripts into beautiful web applications with minimal code. Widely used for ML model demos, data visualization, and internal tools.