Staff SRE: Lead Reliability for Secure Cloud Platforms

Okta

San Francisco (CA)

On-site

USD 210,000 - 260,000

Full time

4 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Okta is seeking an experienced Staff Site Reliability Engineer to join the Emerging Products Group (EPG). You’ll lead reliability across large-scale cloud services, mentor engineers, and drive automation initiatives to reduce toil and improve incident response.

This role emphasizes FedRAMP-aligned security, cross-team collaboration, and platform engineering excellence. You will partner with software and product teams to shape reliability roadmaps, implement SLOs/SLIs, and advance observability

Qualifications

  • Extensive experience architecting and leading the evolution of large-scale production services in AWS and/or GCP.
  • Deep expertise in defining Kubernetes patterns and Linux-based system standards for enterprise-grade production environments.
  • Experience designing multi-region, highly available cloud architectures from the ground up.
  • Experience troubleshooting Kubernetes networking, storage, scheduling, scaling, and workload lifecycle issues.
  • Proven track record of evaluating 'build vs. buy' decisions and setting long-term technical standards for an organization.
  • Extensive experience with Infrastructure as Code technologies such as Terraform and Helm.
  • Strong software engineering skills in Golang and/or Python.
  • Experience building automation and internal engineering platforms.
  • Experience operating and troubleshooting distributed data platforms such as PostgreSQL, Redis, OpenSearch, MySQL, Cassandra, or similar technologies.
  • Strong understanding of cloud networking fundamentals including DNS, load balancing, ingress, TLS, service networking, and traffic management.
  • Strategic experience designing comprehensive observability frameworks and telemetry-driven operational strategies.
  • Experience with or strong interest in AI-assisted engineering and operational automation.

Responsibilities

  • Design, build, and operate large-scale cloud infrastructure and production services.
  • Participate in a global on-call rotation supporting highly available customer-facing systems.
  • Lead incident response efforts and drive post-incident reviews focused on systemic improvements.
  • Define, measure, and improve SLIs, SLOs, and error budgets.
  • Partner with engineering teams to improve service availability, scalability, performance, and resilience.
  • Ensure all infrastructure and operational practices adhere to strict FedRAMP compliance and security mandates.
  • Continuously improve observability through metrics, logging, tracing, dashboards, and alerting.
  • Develop software, automation, and infrastructure using Go, Python, Terraform, and related technologies.
  • Eliminate operational toil through automation, tooling, and platform engineering.
  • Improve deployment safety and operational workflows through CI/CD and GitOps practices.
  • Collaborate on modernizing existing workloads and aligning them with evolving platform capabilities.
  • Build self-service platforms, operational guardrails, and automation that improve developer velocity while maintaining reliability and security.

Skills

AWS/GCP architecture
Kubernetes patterns
Linux enterprise
Multi-region cloud
Kubernetes networking
Terraform
Helm
Golang
Python
Automation & internal platforms
PostgreSQL/Redis/OpenSearch
Cloud networking
Observability/telemetry
AI-assisted engineering interest

Tools

Kubernetes
Terraform
GitOps
OpenTelemetry

Job description

Okta is seeking an experienced Staff Site Reliability Engineer to join the Emerging Products Group (EPG). You’ll lead reliability across large-scale cloud services, mentor engineers, and drive automation initiatives to reduce toil and improve incident response.

This role emphasizes FedRAMP-aligned security, cross-team collaboration, and platform engineering excellence. You will partner with software and product teams to shape reliability roadmaps, implement SLOs/SLIs, and advance observability

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Staff SRE: Cloud Reliability, Automation & Incidents
Senior Staff SRE: Cloud Reliability, Automation & Incidents

Okta • Chicago (IL)

On-site
USD 174,000 - 239,000
Equity
Benefits package
Staff SRE: AI-Driven Cloud Reliability Leader
Staff SRE: AI-Driven Cloud Reliability Leader

Okta • New York (NY)

On-site
USD 174,000 - 239,000
Equity
Bonus
Health, dental and vision insurance
+1
Staff Site Reliability Engineer — Cloud Reliability Leader
Staff Site Reliability Engineer — Cloud Reliability Leader

Okta • Washington

On-site
USD 150,000 - 190,000
Staff SRE: AI Identity Cloud Reliability Leader
Staff SRE: AI Identity Cloud Reliability Leader

Okta • Bellevue (WA)

On-site
USD 180,000 - 240,000
Senior SRE (FedRAMP) – Cloud Reliability & Automation
Senior SRE (FedRAMP) – Cloud Reliability & Automation

Empleora • Northern (KY)

Hybrid
USD 160,000 - 210,000
Senior SRE (FedRAMP) – Cloud Reliability & Security
Senior SRE (FedRAMP) – Cloud Reliability & Security

Visa Hunt • United States

On-site
USD 140,000 - 190,000
Staff SRE: Federal Cloud Reliability & Automation
Staff SRE: Federal Cloud Reliability & Automation

Okta • Chantilly (VA)

On-site
USD 174,000 - 238,000
Staff SRE Federal TS/SCI: Automation & Observability
Staff SRE Federal TS/SCI: Automation & Observability

The available sources do not contain information about the company name for rounx.com. • Washington

Hybrid
USD 174,000 - 238,000
Staff SRE: Federal Networking & Cloud Edge (TS/SCI)
Staff SRE: Federal Networking & Cloud Edge (TS/SCI)

Okta • Chantilly (VA)

On-site
USD 174,000 - 238,000
Equity where applicable
Bonus
Health, dental & vision insurance
+3
Senior Cloud SRE: Automation, Security & Scale
Senior Cloud SRE: Automation, Security & Scale

Okta • San Francisco (CA)

On-site
USD 165,000 - 226,000
Amazing Benefits
Making Social Impact
Fostering Diversity, Equity, Inclusion