Site Reliability Engineer

Kong

United States

On-site

USD 140,000 - 190,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Kong Inc. is seeking a Site Reliability Engineer to strengthen the backbone of our cloud services.

You will build, operate, and automate infrastructure to enable reliability and velocity, working with product engineers to embed best practices throughout the lifecycle. The role requires hands-on cloud, container, and observability expertise, with a focus on incident response, automation, and scalable infrastructure in a dynamic environment.

Qualifications

  • Experience operating production workloads on major cloud providers (AWS, GCP, or Azure).
  • Proficiency in at least one programming or scripting language (Golang, Python, or Bash).
  • Hands-on experience with Docker and Kubernetes for containerized workloads.
  • Knowledge of Infrastructure as Code (Terraform preferred).
  • Familiarity with CI/CD concepts and tools (GitLab CI, Jenkins).
  • Understanding of modern observability stacks (Prometheus, Grafana, ELK).

Responsibilities

  • Build and maintain infrastructure as code to support scalable cloud services.
  • Implement monitoring, logging, and alerting to meet uptime targets.
  • Diagnose production incidents and drive blameless post-mortems.
  • Write automation to reduce toil and enable self-service for engineers.
  • Collaborate with developers to bake reliability into the product lifecycle.
  • Participate in capacity planning, DR drills, and security hardening.
  • Contribute to a fair on-call rotation to ensure platform availability.

Skills

Cloud platforms
Golang
Python
Bash
Observability

Tools

Docker
Kubernetes
Terraform
GitLab CI
Jenkins
Prometheus
Grafana
ELK

Job description

Are you ready to unlock intelligence?

If you don’t think you meet all of the criteria below but are still interested in the job, please apply. Nobody checks every box - we’re looking for candidates that are particularly strong in a few areas, and have some interest and capabilities in others.

The Mission

The Site Reliability Engineering team is the backbone of Kong’s cloud services, responsible for architecting and operating the large-scale infrastructure that powers our customers’ most critical applications. Our mission is to achieve world-class reliability and performance, enabling our product engineering teams to ship features with velocity and confidence. We are the guardians of uptime and the champions of developer delight.

What You’ll Do
  • Build and maintain our core infrastructure as code using tools like Terraform and Ansible.
  • Implement robust monitoring, logging, and alerting systems to ensure our services meet and exceed 99.99% uptime.
  • Resolve production incidents through systematic debugging, and drive the blameless post-mortem process to prevent recurrence.
  • Write automation to reduce operational toil, improve system efficiency, and enable self-service for engineering teams.
  • Collaborate with developers to embed reliability and scalability best practices directly into the application lifecycle.
  • Contribute to our capacity planning, disaster recovery drills, and security hardening processes.
  • Participate in a fair and sustainable on-call rotation to ensure our platform is always available.
What You’ll Bring
  • Experience operating production workloads on a major cloud provider (AWS, GCP, Azure).
  • Proficiency in at least one programming or scripting language, such as Golang, Python, or Bash.
  • Hands-on experience with containerization and orchestration technologies (Docker, Kubernetes).
  • Knowledge of Infrastructure as Code principles and tools (Terraform is a plus).
  • Familiarity with CI/CD concepts and pipeline tools (e.g., GitLab CI, Jenkins).
  • An understanding of modern observability stacks (e.g., Prometheus, Grafana, ELK).

About Kong:

Kong Inc., a leading developer of API and AI connectivity technologies, is building the infrastructure that powers the agentic era. Trusted by the Fortune 500 and startups alike, Kong’s unified API and AI platform, Kong Konnect, enables organizations to secure, manage, accelerate, govern, and monetize the flow of intelligence across APIs and AI models. For more information, visit www.konghq.com.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer, Kong Konnect
Senior Site Reliability Engineer, Kong Konnect

Cacheflow • Washington

On-site
USD 140,000 - 210,000
Senior Site Reliability Engineer, Kong Konnect
Senior Site Reliability Engineer, Kong Konnect

Kong • Washington

On-site
USD 140,000 - 200,000
Datadog observability
Prometheus/Grafana/Thanos stack
ArgoCD-managed CI/CD
Senior SRE, Managed Gateways
Senior SRE, Managed Gateways

Cacheflow • Washington

On-site
USD 190,000 - 230,000
Senior Software Engineering Manager, Managed Gateways SREs
Senior Software Engineering Manager, Managed Gateways SREs

Kong • United States

On-site
USD 128,000 - 165,000
Senior Solutions Engineer
Senior Solutions Engineer

Kong • United States

On-site
USD 100,000 - 130,000
Software Engineer, Core Platform
Software Engineer, Core Platform

Kong Inc. • San Francisco (CA)

On-site
USD 140,000 - 210,000
Staff Technical Support Engineer
Staff Technical Support Engineer

Kong • United States

On-site
USD 110,000 - 170,000
Senior Software Engineer, Konnect Control Plane
Senior Software Engineer, Konnect Control Plane

Kong • United States

On-site
USD 140,000 - 210,000
Senior Software Engineer, Konnect Core Platform
Senior Software Engineer, Konnect Core Platform

Kong • United States

On-site
USD 150,000 - 230,000
Staff Solutions Engineer - New York
Staff Solutions Engineer - New York

Kong • New York (NY)

Hybrid
USD 120,000 - 160,000