Senior Staff SRE: Cloud Reliability, Automation & AI Ops

Okta

San Francisco (CA)

On-site

USD 194,000 - 267,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Equity
Bonus
Health insurance
Dental insurance
Vision insurance
401(k)
Flexible spending account
Paid leave
Parental leave

Job summary

Okta is seeking a Staff Site Reliability Engineer for its Emerging Products Group in the San Francisco Bay area. You will lead reliability initiatives, drive automation, and collaborate with software engineers and product teams to scale secure cloud services.

Expect ownership of on-call, incident reviews, and platform modernization with a bias for automation. The role emphasizes AWS/GCP architecture, Kubernetes patterns, IaC, and proactive reliability engineering across multi-region deployments.

Qualifications

  • Extensive experience architecting and leading large-scale production services in AWS and/or GCP.
  • Deep Kubernetes and Linux system pattern expertise for enterprise-grade production.
  • Design multi-region, highly available cloud architectures from the ground up.
  • Troubleshoot Kubernetes networking, storage, scheduling, and workloads.
  • Build long-term technical standards and evaluate build vs. buy decisions.
  • IaC expertise with Terraform and Helm.
  • Strong Golang and/or Python software engineering skills.
  • Automate and build internal engineering platforms.
  • Operate distributed data platforms (PostgreSQL, Redis, OpenSearch, MySQL, Cassandra).
  • Understand cloud networking (DNS, load balancing, ingress, TLS, service networking).
  • Design observability frameworks and telemetry-driven strategies.
  • Interest or experience in AI-assisted engineering and automation.

Responsibilities

  • Design, build, and operate large-scale cloud infrastructure and production services.
  • Participate in global on-call rotation for highly available customer-facing systems.
  • Lead incident response and post-incident reviews for systemic improvements.
  • Define, measure, and improve SLIs/SLOs and error budgets.
  • Partner with engineering teams to improve availability, scalability, performance and resilience.
  • Ensure infrastructure and practices meet FedRAMP security mandates and audits.
  • Improve observability with metrics, logging, tracing, dashboards, and alerts.
  • Develop automation using Go, Python, Terraform, and related tools.

Skills

Cloud architecture
Kubernetes patterns
Golang/Python
Automation engineering
CI/CD & GitOps
Observability
Security & compliance

Tools

Terraform
Helm
Kubernetes
ArgoCD

Job description

Okta is seeking a Staff Site Reliability Engineer for its Emerging Products Group in the San Francisco Bay area. You will lead reliability initiatives, drive automation, and collaborate with software engineers and product teams to scale secure cloud services.

Expect ownership of on-call, incident reviews, and platform modernization with a bias for automation. The role emphasizes AWS/GCP architecture, Kubernetes patterns, IaC, and proactive reliability engineering across multi-region deployments.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior SRE: AI-Driven Cloud Reliability & Automation
Senior SRE: AI-Driven Cloud Reliability & Automation

Okta • San Francisco (CA)

On-site
USD 165,000 - 227,000
Staff SRE: AI-Driven Cloud Reliability Leader
Staff SRE: AI-Driven Cloud Reliability Leader

Okta • New York (NY)

On-site
USD 174,000 - 239,000
Equity
Bonus
Health, dental and vision insurance
+1
Senior Staff SRE: Cloud Reliability, Automation & Incidents
Senior Staff SRE: Cloud Reliability, Automation & Incidents

Okta • Chicago (IL)

On-site
USD 174,000 - 239,000
Equity
Benefits package
Staff SRE: AI Identity Cloud Reliability Leader
Staff SRE: AI Identity Cloud Reliability Leader

Okta • Bellevue (WA)

On-site
USD 180,000 - 240,000
Staff Site Reliability Engineer — Cloud Reliability Leader
Staff Site Reliability Engineer — Cloud Reliability Leader

Okta • Washington

On-site
USD 150,000 - 190,000
Senior SRE (FedRAMP) – Cloud Reliability & Automation
Senior SRE (FedRAMP) – Cloud Reliability & Automation

Empleora • Northern (KY)

Hybrid
USD 160,000 - 210,000
Senior Cloud SRE: Secure, Scalable Infra & Automation
Senior Cloud SRE: Secure, Scalable Infra & Automation

Okta • Bellevue (WA)

On-site
USD 147,000 - 202,000
Amazing Benefits
Making Social Impact
Fostering Diversity, Equity, Inclusion
Senior Cloud SRE: Automation, Security & Scale
Senior Cloud SRE: Automation, Security & Scale

Okta • San Francisco (CA)

On-site
USD 165,000 - 226,000
Amazing Benefits
Making Social Impact
Fostering Diversity, Equity, Inclusion
Senior SRE Manager: Scale Platforms, Automate & Lead
Senior SRE Manager: Scale Platforms, Automate & Lead

Okta • San Francisco (CA)

On-site
USD 204,000 - 281,000
Health, dental, and vision insurance
401(k) plan
Paid leave including PTO and parental leave
Staff SRE: Federal Cloud Reliability & Automation
Staff SRE: Federal Cloud Reliability & Automation

Okta • Chantilly (VA)

On-site
USD 174,000 - 238,000