Kubernetes SRE Lead: Scalable Cloud Platform on AWS
Okta
San Francisco (CA)
Hybrid
USD 194,000 - 267,000
Full time
14 days+
Application generator
A complete application in a minute — tailored resume and cover letter, ready to send.
Get past ATS filters
Benefits offered by this job
Health, dental, and vision insurance
401(k)
Flexible spending account
Paid leave including PTO and parental leave
Job summary
A leading technology company is seeking a Site Reliability Engineer to build and manage scalable Kubernetes platforms on AWS. The ideal candidate will ensure high availability and performance while optimizing costs. Key qualifications include extensive experience with Kubernetes, AWS, and Terraform, alongside a strong foundation in automation and CI/CD processes. This role is integral to maintaining the reliability and efficiency of cloud-native applications, offering competitive compensation and a vibrant company culture.
Qualifications
4+ years of experience with Kubernetes/Helm.
4+ years of experience with Terraform.
5+ years of experience with AWS.
Experience with multi-region cloud environments.
Strong expertise in Kubernetes platform creation, management, and optimisation.
Responsibilities
Design, implement, and maintain Kubernetes platforms.
Build, manage, and optimize AWS cloud infrastructure.
Utilize Helm to automate application deployments.
Implement and manage Karpenter for scaling.
Automate deployment, scaling, and management of infrastructure.
Skills
Kubernetes/Helm
Terraform
AWS
CI/CD pipelines
Scripting in Python, Bash, or Go
Education
Bachelor's degree in Computer Science, Engineering, or related field
Tools
Prometheus
Grafana
CloudWatch
ELK Stack
Job description
A leading technology company is seeking a Site Reliability Engineer to build and manage scalable Kubernetes platforms on AWS. The ideal candidate will ensure high availability and performance while optimizing costs. Key qualifications include extensive experience with Kubernetes, AWS, and Terraform, alongside a strong foundation in automation and CI/CD processes. This role is integral to maintaining the reliability and efficiency of cloud-native applications, offering competitive compensation and a vibrant company culture.