Git
The winner across all the existing file versioning tools, distributed versioning, fully controllable from the command-line, plenty of configuration and usage options, behind a number of related project that leverage git as backend.
A curated list of Site Reliability and Production Engineering Tools
This page lists names, links and short descriptions. The original list on GitHub is the source and belongs to its authors.
The winner across all the existing file versioning tools, distributed versioning, fully controllable from the command-line, plenty of configuration and usage options, behind a number of related project that leverage git as backend.
Build software better, together. GitHub is the largest code host on the planet with over 13.2 million repositories. Large or small, every repository comes with the same powerful tools. These tools are open to the community for public projects and secure for private projects.
similar to GitHub, GitLab provides git hosting, collaborations, social, automations, and more. GitLab can be both cloud-based and self-hosted using its open-source code.; GitLab includes unlimited free private repositories.; GitLab comes with a continuous integration tool that is more powerful…
Host, manage, and share Git and Mercurial repositories in the cloud. Free, unlimited private repositories for up to 5 developers give teams the flexibility to grow and code without restrictions.
A simple, high-reliability, distributed software configuration management system with these advanced features: project management, built-in web interface, friendly self-hosting, simple networking, all-in-one standalone executable, and much more.
Mercurial is a free, distributed source control management tool. It efficiently handles projects of any size and offers an easy and intuitive interface
Industry standard for large assets (free ≤5 users).See also: Software Reference → Pipeline & Production Management Software
Client-server revision control system. (Source Code) Apache-2.0 C
JIRA is the tracker for teams planning and building great products. Thousands of teams choose JIRA to capture and organize issues, assign work, and follow team activity. At your desk or on the go with the new mobile interface, JIRA helps your team get the job done.
Organize anything, together Trello is the fastest, easiest way to organize anything, from your day-to-day work, to a favorite side project, to your greatest life plans.
is an open-source project management software for cross-functional teams that work agile across both scrum and kanban frameworks.
an online project management software that gives you full visibility and control over your tasks.
Teamwork without email. Asana is our go-to for prioritizing projects, keeping up w/orders & staying on top of a growing to-do list
Easily build, run, and scale your dream workflows on one platform.
ClickUp's #1 rated productivity software is making more productive projects with a beautifully designed and intuitive platform.
The official account for Basecamp®. Helping Basecamp customers every Mon-Fri 9am-6pm CT! Basecamp to help organize the store design, develop fixtures, and manage craftspeople. Primarily through word-of-mouth alone, Basecamp has become the world’s #1 project management tool.
Flexible project management web application. Written using Ruby on Rails framework, it is cross-platform and cross-database.
the most innovative way to manage projects, completely free... forever.
(fka Clubhouse) - Project management for software development teams.
A lesser known feature of GitHub, makes it easy to tie your project management process to your code.
Issue tracking and project management that engineering teams adopt without being told to.
General-purpose bugtracker and testing tool originally developed and used by the Mozilla project. (Source Code) MPL-2.0 Perl
In-app bug and crash reporting with video, logs, network traffic and traces.
Instabug is comprehensive bug reporting and in-app feedback tool for mobile apps. Receive detailed feedback attached with every bug report sent by your testers or users to debug and build better apps, faster.
Bug tracker, fits best for software development. (Demo, Source Code) GPL-2.0 PHP
One of the oldest text editors, free long-standing software project, with a huge amount of functionalities and extensions; implemented and extendable with E-Lisp.
Editor with customizable syntax highlighting, multi-line editing, and plugin support. 🪟 🟢
(Free,Cross-platform,Plugins): electron based editor with numerous plugins and easy modifications. Cross-platform with settings and plugins synchronized through the sync-settings plugin.
is a lightweight but powerful source code editor which runs on your desktop and is available for Windows, macOS and Linux. It comes with built-in support for JavaScript, TypeScript and Node.js and has a rich ecosystem of extensions for other languages (such as C++, C#, Java, Python, PHP, Go) and…
A sophisticated text editor for code, markup and prose. You'll love the slick user interface, extraordinary features and amazing performance.
Vim is a highly configurable text editor built to make creating and changing any kind of text very efficient. (GNU GPL compatible)
A work in progress attempt to improve vim, dropping older/unused OS compatibility, improving the codebase readability, modularity, and maintainability; it has chances to become the next choice of vim users.
Supports plugin-based development workflows.
Easy to use, lightweight text editor; no complex keybindings to remember; the main ones are shown in the main menu.
Editor that brings Apple's approach to operating systems into the world of text editors.
Default text editor for GNOME.
The smartest JavaScript IDE by JetBrains. FREE for Students, check here for more info.
copyright: — Comes bundled with a lot of inspections for Java and Kotlin and includes tools for refactoring, formatting and more.
is the best IDE I've ever used. With PyCharm, you can access the command line, connect to a database, create a virtual environment, and manage your version control system all in one place, saving time by avoiding constantly switching between windows.
Open source workspace server and cloud IDE. (Source Code) EPL-1.0 Docker/Java
Editor targeted towards programmers and web developers (C, JavaScript, Java, PHP, Python and markup languages: HTML, YAML and XML) #c #gtk3.
AI test automation with natural language test creation.
load testing tool for measuring performance of services and distributed systems.
Appium is an open source test automation framework for use with native and hybrid mobile apps. It drives iOS, Android Apps using the WebDriver protocol.
A Chaos Engineering platform (SaaS or On-Prem) with auto discovery features, different attack types, user management and many more.
k6 is a developer centric open source load testing tool for testing the performance of your backend infrastructure. It’s built with Go and JavaScript to integrate well into your development workflow, so you can stay on top of performance without fuzz.
is a tool that makes it fast, easy and reliable testing for anything that runs in a browser.
A small build system with a focus on speed.
An open source build system meant to be both extremely fast, and, even more importantly, as user friendly as possible.
is an open-source, cross-platform family of tools designed to build, test and package software. CMake is used to control the software compilation process using simple platform and compiler independent configuration files, and generate native makefiles and workspaces that can be used in the…
is a tool for automatically generating Makefile.in files compliant with the GNU Coding Standards. Automake requires the use of GNU Autoconf.
Command-line utility which reads a scripted definition of a software project and uses it to generate project files for Visual Studio and GNU Make. Other targets are also being worked on. BSD-3-Clause
Build automation tool mainly for Java. A software project management and comprehension tool. Based on the concept of a project object model (POM), Maven can manage a project's build, reporting and documentation from a central piece of information. (Source Code) Apache-2.0 Java
Automation build tool, similar to make, a library and command-line tool whose mission is to drive processes described in build files as targets and extension points dependent upon each other. (Source Code) Apache-2.0 Java
🛠️ - 🐙- A build tool with a focus on build automation and support for multi-language development.
The most popular automation build tool for many purposes, make is a tool which controls the generation of executables and other non-source files of a program from the program's source files. (Source Code) GPL-3.0 C
Build automation tool similar to Make, written in and extensible in Ruby. (Source Code) MIT Ruby
Nix-based continuous build system
A fast, scalable, multi-language and extensible build system. Used by Google. (Source Code) Apache-2.0 Java
cloud service for software development formerly known as Visual Studio Team Services, Visual Studio Online and Team Foundation Service Preview
Jenkins provides continuous integration services for software development. It is a server-based system that supports SCM tools including AccuRev, CVS, Subversion, Git, Mercurial, Perforce, Clearcase and RTC, and can execute Apache Ant and Apache Maven based projects as well as arbitrary shell…
Bamboo does more than just run builds and tests. It connects issues, commits, test results, and deploys so the whole picture is available to your entire product team – from project managers, to devs and testers, to sys admins.
the previous one of Jenkins
is a continuous integration and continuous delivery platform that helps software teams work smarter, faster.
🛠 - A Java-based build management and continuous integration server from JetBrains.
pipelines build, test, deploy, and monitor your code as part of a single, integrated workflow.
is a hosted continuous integration service used to build and test software projects hosted at GitHub.
automate all aspects of the software development cycle.
Create an Amazing Workflow. Semaphore assumes that your private or open source project is on GitHub. There are no new dependencies, hooks or SSH keys to manage. It works without any change in source code.
Concourse is an open-source continuous thing-doer. Built on the simple mechanics of resources, tasks, and jobs, Concourse presents a general approach to automation that makes it great for CI/CD.
Self-Hosted, Open-Source CI Platform. Based on NodeJS and Docker.
Continuously build, test, release, and monitor apps for every platform.
AppVeyor automates building, testing and deployment of .NET applications.
Continuously test and monitor your APIs after deployments and across environments.
Semi-hosted continuous integration and deployment. Buildkite uses your own infrastructure to run builds so you can test any language or run any deployment scripts. You can run as many parallel agents (and builds) as you want.
Automated Code Review. Continuous Static Analysis designed to complement your unit tests. Similar to CodeClimate.
Automated Code Review. Code Climate hosted software metrics help you ship quality Ruby and JavaScript code faster. Get control of your technical debt with real time static analysis of your code.
Codefresh is a Docker-native CI/CD platform. Instantly build, test and deploy Docker images to Kubernetes
One more cloud based CI service: running tests and deployment
Continuous Integration and deployment for projects written in PHP
Drone is a Continuous Delivery system built on container technology. Drone uses a simple YAML configuration file, a superset of docker-compose, to define and execute Pipelines inside Docker containers.
Comments on style violations in GitHub pull requests. Supports Coffeescript, Go, HAML, JavaScript, Ruby, SCSS and Swift.
Continuous Collaboration - break down the barriers between software developers and the other stakeholders involved in a software development project
Automate and streamline the build-test-release cycle for worry-free, continuous delivery of your product
Provides automated code deployment to EC2 instances.
ElectricFlow/ElectricCommander gives distributed teams shared control and visibility into infrastructure, tool chains and processes. It accelerates and automates the software delivery process to enable agility, predictability and security across many build-test-deploy pipelines
Instantly build and ship code anywhere in one consistent process for your entire team.
Hosted continuous integration and deployment service built on docker
Continuous delivery platform
Declarative, GitOps continuous delivery tool for Kubernetes. (Source Code) Apache-2.0 Go
yen: The best of Git, build & deployment tools combined into one powerful tool that supercharged our development.
GitOps tool with advanced features to build images and deploy them to Kubernetes (integrates with any existing CI system)
Enterprise Kubernetes management platform for deploying applications, databases, Helm charts, and Terraform modules on AWS, GCP, Azure, and Scaleway.
is a tool for building and managing virtual machine environments in a single workflow. With an easy-to-use workflow and focus on automation, Vagrant lowers development environment setup time, increases production parity, and makes the "works on my machine" excuse a relic of the past. It provides…
is an event-driven automation tool and framework to deploy, configure, and manage complex IT systems. It automates common infrastructure administration tasks and ensure that all the components of your infrastructure are operating in a consistent desired state.
is an open source toolkit, written in Python, it is used for configuration management, application deployment, continuous delivery, IT infrastructure automation and automation in general.
infrastructure as code tool for provisioning and managing cloud and on-premises infrastructure.
Provides a file-based interface for provisioning other resources.
Runbook Automation For Modernizing Your Operations.
Alternative to Terraform Cloud/Enterprise. Collaborative Infrastructure Delivery Platform for Terraform :heavy_dollar_sign:
Alternative to Terraform Enterprise with OPA integration, organizational structure, custom hooks, native integrations with other DevOps platforms, and centralized reporting. :heavy_dollar_sign:
Terraform and OpenTofu without the state file bottleneck. Replace the flat state file with a real database. Teams plan in parallel, state is queryable via SQL, and plans run in seconds instead of minutes. :heavy_dollar_sign:
Modern infrastructure as code platform that allows you to use familiar programming languages and tools to build, deploy, and manage cloud infrastructure.
A framework used by platform teams to build the custom platforms tailored to their organisation.
Open-source alternative to Terraform Cloud/Enterprise. GitOps-first and built for scale, security, and reliability across modern VCS providers.
is an open platform for developing, shipping, and running applications. Docker enables you to separate your applications from your infrastructure so you can deliver software quickly working in collaboration with cloud, Linux, and Windows vendors, including Microsoft.
is a daemonless, open source, Linux native tool designed to make it easy to find, run, build, share and deploy applications using Open Containers Initiative (OCI) Containers and Container Images. Podman provides a command line interface (CLI) familiar to anyone who has used the Docker Container…
is a daemon that manages the complete container lifecycle of its host system, from image transfer and storage to container execution and supervision to low-level storage to network attachments and beyond. It is available for Linux and Windows.
is the world's largest library and community for container images Browse over 100,000 container images from software vendors, open-source projects, and the community.
yen: Amazon Elastic Container Registry (ECR) is a fully-managed Docker container registry that makes it easy for developers to store, manage, and deploy Docker container images.
yen: Artifact Repository Manager, can be used as private Docker Registry as well.
is a project that Builds, Stores, and Distributes your Applications and Containers.
is an open-source system for automating deployment, scaling, and management of containerized applications.
is a Docker-native clustering system swarm is a simple tool which controls a cluster of Docker hosts and exposes it as a single "virtual" host.
with Marathon
Provides monitoring for AWS cloud resources and applications, starting with EC2.
Site monitoring based on Lighthouse. See how your scores and metrics changed over time, with a focus on understanding what caused each change. Paid product with a free 30-day trial.
DNS/HTTP/SSL benchmarking with CSV, Excel, PDF, and JSON exports.
is a free software application used for event monitoring and alerting. It records real-time metrics in a time series database (allowing for high dimensionality) built using a HTTP pull model, with flexible queries and real-time alerting.
🛠️ - Stackdriver Logging allows you to store, search, analyze, monitor, and alert on log data and events.
Monitoring tool for ephemeral infrastructure and distributed applications. (Source Code) MIT Go
Error monitoring that helps all software teams discover, triage, and prioritize errors in real-time.
Solve operational problems faster. Loggly helps cloud-centric organizations—organizations that build and manage cloud-facing applications—to solve operational problems faster.
Ship logs from any source, parse them, get the right timestamp, index them, and search them. Logstash is a tool for managing events and logs. You can use it to collect logs, parse them, and store them for later use (like, for searching). Speaking of searching, Logstash comes with a web interface…
一个基于MongoDB的云数据库,提供实时数据同步和API。有免费计划,付费计划起价$0.08/小时。
MongoDB Inc. databases management offer
Application monitoring for all your web apps. It’s about gaining actionable, real-time business insights from the billions of metrics your software is producing, including user click streams, mobile activity, end user experiences and transactions.
Frustration-Free log management. Get started in seconds. Use Papertrail's time-saving log tools, flexible system groups, team-wide access, long-term archives, charts and analytics exports, monitoring webhooks, and 45-second setup.
Free all-in-one website health scanner. Core Web Vitals, SEO, WCAG 2.1 accessibility, and best practices. AI-generated action plan. No signup required.
Test the load time of that page, analyze it, and find bottlenecks.
Premium hosted website and server monitoring tool. All your activity syncs in real time - from starting new instances to upgrading or deleting old ones. Work wherever you want - through web, mobile, API or directly with your provider. Everything stays in sync.
Enterprise-class software for monitoring of networks and applications. (Source Code) GPL-2.0 C
Better monitoring for your Rails applications. Get detailled statistics on your site's performance with mean and 90th percentile measurements.
Centralized dashboard tracking real-time status and outages for 1,000+ popular APIs and services (AWS, Stripe, GitHub, Twilio, etc.). Monitor third-party dependencies, get instant outage alerts, reduce MTTR.
is an analytics platform that enables you to query and visualize data, then create and share dashboards based on your visualizations. Easily visualize metrics, logs, and traces from multiple sources such as Prometheus, Loki, Elasticsearch, InfluxDB, Postgres, Fluentd, Fluentbit, Logstash and many…
fast, resource-effective and scalable open source time series database. May be used as long-term remote storage for Prometheus. Supports PromQL.
Detects cloud waste and helps DevOps/platform teams identify quick cloud cost optimization opportunities.
is a set of components that can be composed into a highly available metric system with unlimited storage capacity, which can be added seamlessly on top of existing Prometheus deployments.
Uptime monitoring & Statuspages
Open-source SSL/TLS certificate expiry monitoring tool with email alerts
Open-source DNS propagation monitoring tool with global DNS server coverage
AI-powered outage aggregator tracking 100+ cloud services with Telegram alerts
Real-time status dashboard for 75+ AI services (OpenAI, Anthropic, Gemini, Mistral, etc.) with starring, email/webhook alerts, 30-day uptime history, and embeddable SVG badges.
Universal SQL interface to any cloud API
Uptime monitoring, incident management, and status pages.
Instantly diagnose slowdowns and anomalies in your infrastructure.
Datadog is a monitoring service for IT, Operations and Development teams who write and run applications at scale, and want to turn the massive amounts of data produced by their apps, tools and services into actionable insight.
Developer-first uptime monitoring with HTTP, DNS, TCP, ICMP, and heartbeat checks, dependency intelligence for 80+ providers, hosted status pages, incident management, and a full developer surface (CLI, SDKs, Terraform provider, MCP server).
Listen for pings and sends alerts when pings are late. (Source Code) BSD-3-Clause Python/Docker
Uptime monitoring for websites, APIs, and cron jobs, with integrated status pages.
Uptime monitoring with 30-second checks on free tier, consecutive-check alert confirmation to cut false positives, hosted status pages, and a built-in MCP server for AI agents.
Code-Native Data Privacy - embed privacy controls in your application code to detect and monitor PII.
OpenTelemetry Native Observability, built on CNCF Open Standards such as PromQL, Perses and OTLP with full cost control. Supporting Metrics, Traces and Logs with full custom dashboarding and alerting capabilities.
AI DevOps monitoring platform by monitoring your CI workflows, detect anomalies, and provide actionable fixes.
A Full-Stack Cloud Observability Platform designed to empower developers and organizations to monitor, optimize, and streamline their applications and infrastructure in real-time.
Boost GitHub Actions speed by 2x and cut costs by up to 75%, with smarter caching, deep CI insights, and zero-config setup.
eBPF-based GPU causal observability agent. Traces CUDA APIs and host kernel events to build causal chains explaining GPU latency. Includes MCP server for AI-assisted incident investigation.
AWS security auditing CLI that runs 17 checks across IAM, S3, EC2, VPC, and RDS with built-in remediation engine generating AWS CLI commands and Terraform snippets.
Uptime, content, and dependency monitoring with multi-region verification, status pages, and incident management.
Shockingly good uptime monitoring, alerts, incident management, and status pages.
Lightweight columnar log analytics database for SRE workflows, with a pipe-style query language inspired by SPL for investigating production logs.
Open-source multi-cluster Kubernetes dashboard with AI-powered operations, MCP server bridging kubeconfig to LLM agents, and real-time observability across edge and cloud clusters. CNCF Sandbox.
API monitoring, analytics, and request logging for REST APIs, with lightweight open-source SDKs for Python, Node.js, Go, .NET, and Java.
Cross-repo infrastructure dependency discovery and change impact analysis for multi-repo environments using Terraform, Docker, Helm, and more.
HTTP monitoring with TCP kernel telemetry, 6-phase latency breakdown, Server-Timing header capture, Cloudflare CDN enrichment, and built-in incident management with on-call scheduling.
Real-time AI agent monitoring dashboard for OpenClaw agents. Track Gateway status, sessions, token usage & trends.
TUI observability for AI coding agents. Track cost, tokens, tool failures, latency, anomalies, health, diffs, and CI gates across Claude Code, Codex CLI, Gemini CLI, Aider, and Cursor exports.
Crypto-paid uptime monitoring for side projects. Pay per monitor with USDC on Base; webhook alerts on down/up state changes.
Cron, heartbeat, and HTTP uptime monitoring for background jobs and services, with concurrent-job (run_id) correlation, duration/hang alerts, LOG pings for mid-run progress, incident management, and status pages. One curl ping instruments a job; no agent or SDK. Free tier: 50 monitors, 200K…
Continuous monitoring of blockchain RPC providers, bridges and oracles. Multi-region Prometheus probes of latency, tx-landing success and finality with public dashboards and open methodology.
Observability platform for LLM and AI agent applications, with tracing, evals, prompt management, and a gateway across 250+ models.
Deterministic CI failure analysis CLI that classifies build logs into explainable failure types with evidence and fix steps.
Monitoring for uptime, performance, broken links, SSL certificates, and DNS, with hosted status pages.
OpenTelemetry-native synthetic monitoring with HTTP and Playwright browser checks, monitoring-as-code via YAML and CLI, and enriched OTLP export to any OTel backend.
On-Call & SRE platform with StatusPages built-in
AI-powered incident management and response automation.
Incident Management to Support the DevOps Lifecycle, from First Alert to Post-Incident Review
We make alerts work for you. We provide the tools you need to design meaningful, actionable alerts and ensure the right people are notified.
Now FireHydrant
incident response platform with status pages built-in
Open-source CLI to manage Rootly incidents, alerts, services, teams, and on-call schedules from the terminal.
Simple alerting tool, with declarative syntax and builtin providers.
Uptime monitoring, incident management, and status pages.
Incident management, on-call alerting, and public, customer, and private status pages.
Investigate Prometheus alerts, Jira/Pagerduty/Opsgenie tickets automatically using AI.
Open-source AI on-call developer
Debug Production x10 faster with AI.
Reliability Shift Left platform. Generate dashboards, alerts, SLOs from YAML. Verify metrics exist before deploy. Block deploys when error budget exhausted.
Incident management platform with on-call scheduling, real-time collaboration, and automated escalations.
Shared causal traces for incident response. Captures pre-alert causal chains across services and assembles them into a shared, replayable artifact before the war room starts—open-source SDKs.
Open-source, self-hosted incident management with alert ingestion, on-call scheduling, escalation policies, AI-powered post-mortems, and Slack/Teams integration. AGPLv3 — self-hosted alternative to PagerDuty and Grafana OnCall.
Orchestrator across IT, HR, CRM.
Relationships between businesses and their customers can be hard. Better customer service starts with better communication Zendesk brings all your customer conversations into one place.
communicate incidents and maintenance effectively with a beautiful hosted status page.
Uptime monitoring & Statuspages
AI-powered outage aggregator tracking 100+ cloud services with Telegram alerts
Quick and beautiful status page.
A single-site, alternative to https://statuspage.io written in PHP with the Laravel project, supporting both SQLite and MySQL databases.
A platform for building no-code, holistic, Internal Developer Portals.
An open platform for building developer portals.
eBPF-based GPU causal observability agent. Traces CUDA APIs and host kernel events to build causal chains explaining GPU latency. Includes MCP server for AI-assisted incident investigation.
(open source)
MCP server with 52 tools for managing Tailscale tailnets from AI assistants like Claude Code and Cursor.
AI-powered multi-cluster Kubernetes management console with MCP server (kc-agent) for AI-assisted cluster operations, pod inspection, deployment management, and real-time observability across distributed environments.
Deep research agent for your infra - sandboxed, read-only, covers AWS, GCP, Azure, Kubernetes, GitHub and GitLab.
Open source (Apache 2.0) AI SRE agent that autonomously investigates incidents and performs root cause analysis across AWS, Azure, GCP, and Kubernetes. Self-hosted via Docker Compose or Helm, works with major LLM providers or local models via Ollama.
AI SRE built on a versioned resource graph of your infrastructure, for root cause analysis and predicting the impact of changes before they ship.
Open source Kubernetes visibility tool with a built-in MCP server for AI-assisted cluster operations — topology, service traffic, events, logs, and a 31-check best-practices audit.
AI-native ops agent that gives agents production-safe execution with human review and a built-in knowledge graph.
Unified AI agentic platform for cloud ops, offering AI SRE, AI FinOps, AI Kubernetes Ops, and AI CloudOps assistants that automate alert triage, root-cause analysis, and cost optimization.
Self-hosted AI SRE agent that goes beyond on-call incident resolution.
awesome-foss/awesome-sysadmin
A curated list of amazingly awesome open-source sysadmin resources.
unixorn/awesome-zsh-plugins
A collection of ZSH frameworks, plugins, themes and tutorials.
agarrharr/awesome-cli-apps
🖥 📊 🕹 🛠 A curated list of command line apps
shuaibiyy/awesome-tf
Curated list of resources on HashiCorp's Terraform and OpenTofu
wmariuss/awesome-devops
A curated list of awesome DevOps platforms, tools, practices and resources
rootsongjc/awesome-cloud-native
A curated list for awesome cloud native tools, software and tutorials.