About The Role
The Site Reliability Engineer will design, operate, and improve the production infrastructure supporting high-volume, customer-facing services. The role focuses on Kubernetes, AWS, observability, incident response, and automation across distributed systems where availability, latency, and operational safety are critical.
About The Role
The Site Reliability Engineer will design, operate, and improve the production infrastructure supporting high-volume, customer-facing services. The role focuses on Kubernetes, AWS, observability, incident response, and automation across distributed systems where availability, latency, and operational safety are critical.
The engineer will partner with software teams to define service-level objectives, eliminate recurring failure modes, and build reliable deployment and recovery workflows. This role has direct ownership of production health and meaningful influence over platform architecture, engineering standards, and operational practices.
Key Responsibilities
- Operate and improve highly available production services running on AWS and Kubernetes, including capacity planning, scaling, and failure recovery
- Define and maintain SLIs, SLOs, error budgets, and operational dashboards using tools such as Prometheus, Grafana, and OpenTelemetry
- Automate infrastructure provisioning and configuration with Terraform, Helm, and GitHub Actions or equivalent CI/CD systems
- Lead incident response, coordinate technical mitigation, and produce clear post-incident reviews with measurable corrective actions
- Harden deployment pipelines with progressive delivery, automated rollback, health checks, and change-management controls
- Identify and eliminate recurring toil through Python or Go automation, platform tooling, and self-service workflows
- Collaborate with application engineers on performance tuning, resilience testing, dependency management, and production readiness reviews
What We Are Looking For
- 3-8 years of experience in site reliability engineering, DevOps, platform engineering, or a closely related production infrastructure role
- Hands-on experience operating Kubernetes workloads in production, including deployments, networking, storage, ingress, and troubleshooting
- Strong knowledge of AWS services such as EC2, EKS, IAM, VPC, RDS, S3, and CloudWatch
- Proficiency with infrastructure as code and delivery tooling, including Terraform, Helm, Git, and CI/CD pipelines
- Experience building observability systems with metrics, logs, traces, alerting, and on-call practices using tools such as Prometheus, Grafana, Datadog, or OpenTelemetry
- Proficiency in Python, Go, or a similar programming language, with a track record of replacing manual operational work with reliable automation
- Bachelor's degree in computer science, engineering, or a related technical field, or equivalent practical experience
- Bonus: experience with service meshes, distributed systems, chaos engineering, compliance-focused infrastructure, or multi-region production environments