Senior Cloud Reliability Engineer (Kubernetes & Go)

Upbound - Job Posting

Chicago (IL)

On-site

USD 150,000 - 210,000

Full time

10 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Upbound is seeking an experienced engineer to run production cloud services at scale, focusing on reliability, observability, and incident response. You will operate Kubernetes controllers, implement monitoring with Prometheus, Grafana, and OpenTelemetry, and collaborate closely with customers to reproduce and fix complex issues.

In this remote role, you will own high‑priority features, contribute to Crossplane open‑source work, and help standardize runbooks and post‑incident reviews.

Qualifications

  • Proven experience running production cloud services at scale.
  • Strong troubleshooting ability in distributed systems with observability tooling.
  • Hands-on experience with Prometheus, Grafana, OpenTelemetry and distributed tracing.
  • Experience building and operating Kubernetes controllers with API server interactions.
  • Comfort working directly with customers to reproduce and fix complex issues.
  • Demonstrated ownership for solving problems during incidents affecting workloads.
  • Commitment to high-quality work and improving operational resilience.
  • Customer empathy and focus on debuggability as a product feature.
  • Clear communication and collaboration to support customers and team.

Responsibilities

  • Build and operate Upbound Spaces in production across multi-tenant SaaS environments.
  • Own high-priority features requested by customers and deliver new capabilities.
  • Investigate and debug complex customer-side issues, including multi-control-plane scenarios.
  • Produce thoughtful design documents for new initiatives and post-incident reviews.
  • Support the full lifecycle of scalable cloud services from discovery to operations.
  • Write Go code interfacing with the Kubernetes API with emphasis on observability and reliability.
  • Deploy, manage, and troubleshoot Kubernetes services in production using metrics, logs, and traces.
  • Build and maintain tooling for customer debugging, control plane health analysis, and incident-response automation.
  • Create documentation, runbooks, and blog posts to support shipped features.
  • Support release processes for self-hosted Spaces distributions and diagnose issues in customer environments.
  • Participate in an on-call rotation for Upbound Cloud and resolve incidents end-to-end.

Skills

Production cloud ops
Troubleshooting distributed systems
Observability tools
Kubernetes controllers
Customer focus
Ownership & accountability
Quality & reliability
Communication & collaboration
Team-oriented mindset

Tools

Kubernetes
Prometheus
Grafana
OpenTelemetry

Job description

Upbound is seeking an experienced engineer to run production cloud services at scale, focusing on reliability, observability, and incident response. You will operate Kubernetes controllers, implement monitoring with Prometheus, Grafana, and OpenTelemetry, and collaborate closely with customers to reproduce and fix complex issues.

In this remote role, you will own high‑priority features, contribute to Crossplane open‑source work, and help standardize runbooks and post‑incident reviews.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Software Engineer – Cloud Platform & Operations
Senior Software Engineer – Cloud Platform & Operations

TalentCloud Group • Indiana

Hybrid
USD 100,000 - 130,000
Competitive salary
Equity
Flexible working hours
+1
Senior Platform Reliability Engineer - Kubernetes & Observability
Senior Platform Reliability Engineer - Kubernetes & Observability

PLP Group • New York (NY)

On-site
USD 75,000 - 130,000
Remote Senior Staff SRE — Kubernetes, CI/CD & Cloud Scale
Remote Senior Staff SRE — Kubernetes, CI/CD & Cloud Scale

Far Coder • Northern (KY)

Hybrid
USD 170,000 - 227,000
Senior Platform & Reliability Engineer — Kubernetes & Go
Senior Platform & Reliability Engineer — Kubernetes & Go

United States Digital Space LLC • United States

Remote
USD 150,000 - 210,000
Remote-first team
Flexible working hours
Async-friendly culture
+2
Senior Cloud Infra Engineer (Go) — Remote, Equity & PTO
Senior Cloud Infra Engineer (Go) — Remote, Equity & PTO

RunPod Inc. • United States

On-site
USD 150,000 - 200,000
Equity
Generous benefits
Flexible PTO
+2
Senior Backend Engineer - Kubernetes & Cloud (Remote)
Senior Backend Engineer - Kubernetes & Cloud (Remote)

Red Hat, Inc. • Raleigh (NC)

Hybrid
USD 157,248 - 195,680
Senior Software Engineer, Kubernetes Spaces Platform
Senior Software Engineer, Kubernetes Spaces Platform

Upbound • Chicago (IL)

On-site
USD 120,000 - 160,000
Kubernetes Engineer
Kubernetes Engineer

Veriipro • Phoenix (AZ)

On-site
USD 120,000 - 170,000
Health insurance
401(k) plan
Paid time off
Go Cloud Infrastructure Engineer - Production Systems
Go Cloud Infrastructure Engineer - Production Systems

Thunder Compute • San Francisco (CA)

On-site
USD 100,000 - 140,000
Senior Escalation Engineer: Kubernetes & Go Specialist
Senior Escalation Engineer: Kubernetes & Go Specialist

Everpure • Georgia

On-site
USD 107,000 - 230,000
Flexible time off
Wellness resources
Company-sponsored team events