Staff Site Reliability Engineer (Kubernetes)

Okta

Washington

On-site

USD 180,000 - 240,000

Full time

2 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Work from home opportunities
Health + Wellness
Financial Benefits
Pay + Incentives
Time Off
Everyday Living
Resources

Job summary

Okta is seeking a Site Reliability Engineer to build and manage Kubernetes platforms on AWS, focusing on reliability, scalability, and security. You will design highly available clusters, automate deployments, and optimize costs while supporting cloud-native applications.

The role requires hands-on experience with Kubernetes, Helm, Karpenter, Istio, and AWS services, plus strong scripting (Python/Bash/Go) and CI/CD tooling. Travel may be required for onboarding at SF or Chicago offices.

Qualifications

  • Bachelor's degree in CS, Engineering or related field is required or equivalent experience.
  • Hands-on experience with Kubernetes platforms and cloud-native architectures.
  • Experience with AWS cloud services and multi-region environments.
  • Proficiency in Kubernetes tooling: Helm, Karpenter, Istio, and related tech.

Responsibilities

  • Design, implement, and maintain highly available, scalable Kubernetes platforms.
  • Manage AWS infrastructure (EKS, RDS, S3, VPCs, IAM) with cost and security best practices.
  • Develop Helm charts for production deployments and automate workflows with CI/CD.
  • Implement and manage Karpenter for dynamic cluster scaling.
  • Configure Istio for service mesh, observability, and secure communications.
  • Automate platform deployment, scaling, and CI/CD integration with minimal downtime.
  • Respond to incidents, troubleshoot performance, availability, and security issues.
  • Ensure security and compliance across cloud platforms and Kubernetes clusters.
  • Document procedures and promote knowledge sharing across teams.

Skills

Kubernetes
AWS
Istio
Karpenter
Helm
Terraform
CI/CD
Python
Automation
Security

Education

Bachelor's degree in CS/Engineering or related field
Certifications: CKA/CKAD/AWS DevOps Engineer (preferred)

Tools

Helm
Karpenter
Jenkins
GitLab
CircleCI
Terraform
Ansible
Spinnaker

Job description

  • The Site Reliability Engineer (SRE) will play a key role in building and managing Kubernetes platforms that support cloud-native applications and services
  • This position focuses on architecting and managing reliable, scalable, and secure Kubernetes-based platforms on AWS, ensuring high availability and performance while optimizing costs and automation. The ideal candidate will have hands-on experience with AWS infrastructure, Kubernetes platform creation, Helm charts, Karpenter scaling, and Istio service mesh
  • Kubernetes Platform Creation: Design, implement, and maintain highly available, scalable, and fault-tolerant Kubernetes platforms. Ensure clusters are optimized for production workloads, providing high resilience and operational efficiency
  • AWS Infrastructure Management: Build, manage, and optimize AWS cloud infrastructure, including EKS,ECS, S3, VPCs, RDS, IAM, and more. Implement best practices for cost management, scaling, and security within AWS
  • Helm Management: Utilize Helm to automate and streamline the deployment of applications and services to Kubernetes clusters. Create, maintain, and manage Helm charts for production-ready deployments
  • Karpenter Implementation: Implement and manage Karpenter to dynamically scale Kubernetes clusters in response to workload demands
  • Istio Service Mesh Management: Configure and manage Istio to provide service-to-service communication, security, and observability within the Kubernetes clusters. Enable fine-grained traffic management, service discovery, and policy enforcement
  • Platform Automation & Scaling: Automate the deployment, scaling, and management of infrastructure and applications. Work with CI/CD pipelines to ensure a seamless flow from development to production with minimal downtime
  • Incident Management & Troubleshooting: Respond to incidents, troubleshoot, and resolve system issues related to performance, availability, and security in a timely and effective manner
  • Security & Compliance: Design and implement secure cloud infrastructure with appropriate access controls, network security, and compliance frameworks
  • Documentation & Knowledge Sharing: Create and maintain detailed documentation for Kubernetes platform setup, operational procedures, and best practices. Promote knowledge sharing across teams
Benefits
  • Work from home opportunities
  • Health + Wellness
  • Financial Benefits
  • Pay + Incentives
  • Time Off
  • Everyday Living
  • Resources
  • 4+ years of Experience with Terraform
  • Hands-on experience with Helm for Kubernetes application deployment and management
  • Proficiency in CI/CD pipelines and automation tools (e.g., Jenkins, GitLab, CircleCI, Terraform, Ansible, Spinnaker)
  • Experience with multi-region cloud environments
  • 5+ years of Experience with AWS
  • Expertise in managing and securing Istio for service mesh, including traffic management, security, and observability features
  • Experience with monitoring, logging, and alerting tools such as Prometheus, Grafana, CloudWatch, and ELK Stack
  • Proven experience with AWS (EC2, RDS, S3, CloudFormation, IAM, etc.) and solid understanding of cloud-native architectures
  • Strong expertise in Kubernetes platform creation, management, and optimisation (e.g., setting up highly available clusters, networking, and storage)
  • 4+ years of experience with Kubernetes/Helm
  • Strong scripting and automation skills in Python, Bash, or Go for infrastructure management and platform automation
  • Practical experience with Karpenter for dynamic scaling of Kubernetes clusters and optimising resource usage
  • Requires in-person onboarding and travel to our San Francisco, CA HQ office or our Chicago office during the first week of employment
  • This position requires the ability to access federal environments and/or have access to protected federal data. As a condition of employment for this position, the successful candidate must be able to submit documentation establishing U.S. Person status (e.g. a U.S. Citizen, National, Lawful Permanent Resident, Refugee, or Asylee. 22 CFR 120.15) upon hire
  • Understanding of security best practices for cloud platforms and Kubernetes (e.g., role-based access control (RBAC), encryption, and compliance frameworks)
  • Familiarity with Docker and containerization principles
  • Bachelor's degree in Computer Science, Engineering, or related field (or equivalent professional experience)
  • Certifications (Preferred): CKA (Certified Kubernetes Administrator), CKAD (Certified Kubernetes Application Developer), or AWS Certified DevOps Engineer are highly desirable
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff Site Reliability Engineer - Kubernetes
Staff Site Reliability Engineer - Kubernetes

Talanto • Bellevue (WA), Northern (KY)

Hybrid
USD 174,000 - 267,000
Staff Site Reliability Engineer - Kubernetes
Staff Site Reliability Engineer - Kubernetes

Okta • New York (NY)

On-site
USD 140,000 - 190,000
Staff Site Reliability Engineer - Kubernetes
Staff Site Reliability Engineer - Kubernetes

Segment (Twilio) • Chicago (IL), New York (NY), Bellevue (WA)

On-site
USD 174,000 - 214,000
Equity
Health insurance
Paid leave
+2
Staff Site Reliability Engineer - Kubernetes
Staff Site Reliability Engineer - Kubernetes

Triwill Group • Washington (IL)

On-site
USD 194,000 - 267,000
Equity
Bonus
Health, dental & vision insurance
+3
Kubernetes Platform SRE — AWS Cloud Automation Lead
Kubernetes Platform SRE — AWS Cloud Automation Lead

Triwill Group • Washington (IL)

On-site
USD 194,000 - 267,000
Equity
Bonus
Health, dental & vision insurance
+3
Kubernetes Software Integration Engineer
Kubernetes Software Integration Engineer

Gigatec • Annapolis (MD)

On-site
USD 90,000 - 130,000
100% Paid Healthcare
10% 401k in every paycheck
100% Fully Vested
Staff Site Reliability Engineer - Kubernetes
Staff Site Reliability Engineer - Kubernetes

Okta • Washington

On-site
USD 194,000 - 267,000
Equity
Bonus
Health insurance
+5
Principal Security Engineer
Principal Security Engineer

Mass Digital Health • Boston (MA)

On-site
USD 120,000 - 150,000
Site Reliability Engineer
Site Reliability Engineer

Flanksource Inc. • Mission (KS)

Remote
USD 90,000 - 130,000
100% remote work
Flexible hours
Opportunity to work with cutting-edge technology
Kubernetes Platform SRE | AWS, Istio & Automation
Kubernetes Platform SRE | AWS, Istio & Automation

Okta • Washington

On-site
USD 194,000 - 267,000
Equity
Bonus
Health insurance
+5