About The Role
The role owns the infrastructure that keeps production systems online - CI/CD pipelines, Kubernetes clusters, cloud networking, and the observability stack that detects problems before customers do.
The engineer will work closely with backend and platform teams to automate everything: deployments, scaling, incident response, and infrastructure provisioning across a multi-service, high-traffic environment.
Key Responsibilities
- Build and maintain CI/CD pipelines (GitHub Actions, GitLab CI, or CircleCI) enabling safe, frequent deployments across dozens of microservices
- Operate and scale Kubernetes clusters on AWS or GCP, including cluster upgrades, autoscaling policies, and workload optimization
- Define infrastructure as code using Terraform and Terragrunt, with peer-reviewed modules and automated plan/apply workflows
- Design and maintain observability tooling - Prometheus, Grafana, and distributed tracing - with actionable SLOs and alerting that minimizes noise
- Lead incident response and blameless postmortems; drive reduction of MTTR through runbooks, automation, and chaos testing
- Harden production environments: IAM policies, network segmentation, secrets management, and security patching pipelines
- Partner with development teams to improve service reliability, defining error budgets and capacity plans as the platform scales
What We Are Looking For
- 3+ years of experience in DevOps, SRE, or infrastructure engineering, with production ownership of systems serving significant traffic
- Deep hands-on experience with Kubernetes: deployment, networking, Helm, and troubleshooting in production environments
- Strong Terraform/IaC skills and proficiency with at least one major cloud provider (AWS, GCP, or Azure)
- Solid scripting and automation skills in Python, Bash, or Go
- Experience building observability from the ground up: metrics, logs, traces, and alert design (Prometheus, Datadog, or similar)
- Bachelor's degree in Computer Science or equivalent practical experience
- Bonus: Experience with service mesh (Istio/Linkerd), GitOps tooling (ArgoCD/Flux), multi-region failover design, or compliance environments (SOC 2, PCI)