About The Role
The Site Reliability Engineer owns the systems and practices that keep production services available, performant, and predictable at scale. The role spans Kubernetes-based infrastructure, cloud networking, observability, incident response, and automated delivery across distributed systems.
About The Role
The Site Reliability Engineer owns the systems and practices that keep production services available, performant, and predictable at scale. The role spans Kubernetes-based infrastructure, cloud networking, observability, incident response, and automated delivery across distributed systems.
You will partner with software engineers and platform teams to reduce operational risk, improve deployment safety, and eliminate recurring sources of toil. The work directly impacts uptime, latency, recovery time, and the ability to ship changes without compromising customer experience.
Key Responsibilities
- Design and operate highly available production infrastructure on AWS, GCP, or Azure using Kubernetes, Terraform, and managed data services
- Build and maintain observability systems with Prometheus, Grafana, OpenTelemetry, and centralized logging to track service health, latency, capacity, and error budgets
- Automate deployment, scaling, and recovery workflows through CI/CD pipelines, infrastructure as code, and self-service platform tooling
- Lead incident response for production outages, coordinate mitigation, and document clear post-incident analyses with measurable follow-up actions
- Define and enforce SLOs, SLIs, alerting standards, and reliability practices across critical services
- Improve system performance and resilience through load testing, capacity planning, failure-mode analysis, and controlled disaster-recovery exercises
- Collaborate with application engineers on architecture reviews, operational readiness, and safe rollout strategies including canary and blue-green deployments
What We Are Looking For
- 3–8 years of experience in site reliability engineering, DevOps, production engineering, or a closely related infrastructure role
- Hands-on experience operating Kubernetes and containerized workloads in production, including cluster security, networking, upgrades, and capacity management
- Strong proficiency with at least one cloud platform, preferably AWS, GCP, or Azure, and infrastructure as code using Terraform or an equivalent tool
- Experience building production observability with metrics, logs, and traces using tools such as Prometheus, Grafana, OpenTelemetry, Datadog, or equivalent
- Solid programming or scripting ability in Go, Python, or Bash, with a focus on automation, testing, and maintainable operational tooling
- Practical experience with incident management, SLOs, error budgets, on-call operations, and root-cause analysis for distributed systems
- Bachelor’s degree in computer science, engineering, or a related technical field, or equivalent professional experience; Bonus: experience with service meshes, Argo CD, Kafka, multi-region architectures, chaos engineering, or compliance-focused infrastructure