Site Reliability Engineer

Okta

United States

On-site

USD 120,000 - 180,000

Full time

4 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

The World’s Identity Company is seeking a Site Reliability Engineer II to join the Emerging Products Group. You will help build reliable, scalable cloud services and automation, collaborating with software engineers and senior SREs to ensure production systems run smoothly.

You will implement observability, CI/CD, and GitOps practices, automate repetitive tasks, and contribute to incident response and post‑mortem analysis to drive continuous reliability improvements.

Qualifications

  • Hands-on experience operating production services in cloud environments.
  • Strong understanding of SRE concepts including SLOs, SLAs and post‑mortems.
  • Proficiency with automation via IaC and modern CI/CD pipelines.

Responsibilities

  • Operate large-scale cloud infrastructure and production services.
  • Participate in on-call rotation and incident response.
  • Monitor SLIs/SLOs and drive improvements to reliability and performance.

Skills

Kubernetes
Terraform
Go
Python
CI/CD
Observability
AWS/GCP
SRE concepts
On-call

Tools

Datadog
Splunk
OpenSearch
PostgreSQL
Redis

Job description

**Secure Every Identity, from AI to Human


**Identity is the key to unlocking the potential of AI. the company secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence.


This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk.


Get to Know the company

the company is The World’s Identity Company. We free everyone to safely use any technology—anywhere, on any device or app. Our Workforce and Customer Identity Clouds enable secure yet flexible access, authentication, and automation that transforms how people move through the digital world, putting Identity at the heart of business security and growth.


At the company, we celebrate a variety of perspectives and experiences. We are not looking for someone who checks every single box, we’re looking for lifelong learners and people who can make us better with their unique experiences.


Join our team! We’re building a world where Identity belongs to you.


The Engineering Opportunity

We are looking for a Site Reliability Engineer II to join the company’s Emerging Products Group (EPG). Our mission is to build highly reliable, scalable, and secure cloud services that our customers can trust. We embrace an automation-first mindset and continuously invest in platform engineering, observability, and operational excellence to enable our engineering teams to move quickly and safely.


This role is ideal for a Site Reliability Engineer who is passionate about solving technical challenges, building automation, and maintaining the reliability of production systems. You will serve as a key technical contributor within the EPG SRE organization, collaborating closely with software engineers and senior SREs to support and improve world-class cloud services.


The ideal candidate aligns with the philosophy of “if you have to do it more than once, automate it” and possesses a strong drive for continuous learning, operational stability, and software engineering.


What You’ll Be DoingReliability & Operations

  • Support and Maintain: Assist in operating large-scale cloud infrastructure and production services, focusing on deployment, maintenance, and day-to-day health.
  • On-Call Rotation: Participate in an SRE on-call rotation supporting highly available customer-facing systems.
  • Incident Response: Contribute to incident response efforts, assist in troubleshooting, and participate in post-incident reviews to identify systemic fixes.
  • Service Metrics: Monitor and help maintain Service Level Indicators (SLIs) and Service Level Objectives (SLOs), raising awareness when error budgets are threatened.
  • Observability: Implement and configure observability configurations using metrics, logging, tracing, dashboards, and alerting patterns established by the team.

Engineering & Automation

  • Automation Development: Develop software, tools, and infrastructure-as-code updates using Go, Python, Terraform, and related technologies.
  • Toil Reduction: Identify opportunities to eliminate operational toil through automation and self-service tooling.
  • CI/CD & GitOps: Utilize and maintain CI/CD pipelines and GitOps workflows to ensure safe, repeatable deployment of infrastructure and services.
  • Platform Alignment: Collaborate with engineering teams to migrate and integrate existing workloads into modern platform capabilities.

Collaboration & Execution

  • Project Execution: Execute assigned technical tasks and sub-projects from design through to production rollout with guidance from senior SREs.
  • Knowledge Sharing: Document operational procedures, runbooks, and share technical knowledge within the team.
  • Standard Adherence: Actively follow and adopt operational best practices, security standards, and reliability principles.

Innovation

  • Modern Tooling: Learn and apply modern operational techniques—including AI-assisted tools—to increase daily productivity, accelerate troubleshooting, and simplify automation workflows.

Our Tech StackCategoryTechnologiesInfrastructure/Orchestration

Kubernetes (EKS/GKE), Terraform, Helm, Git, ArgoCD, GitOps


Programming

Golang, Python


Observability

Datadog, Splunk


Data Stores

PostgreSQL, Redis, OpenSearch


What We Are Looking ForTechnical Excellence

  • Cloud Operations: Hands-on experience operating production services in AWS and/or GCP environments.
  • Kubernetes Foundation: Solid, practical understanding of Kubernetes, container runtimes, and troubleshooting application workloads.
  • Infrastructure as Code: Experience using Terraform and Helm to provision and manage cloud resources.
  • Programming Skills: Proficiency in software development using Golang and/or Python.
  • Networking Basics: Good understanding of cloud networking fundamentals, including DNS, load balancing, HTTPS/TLS, and basic traffic routing.
  • Data Platforms: Familiarity operating and querying distributed data stores (e.g., PostgreSQL, Redis, OpenSearch).
  • Observability: Hands-on experience configuring metrics, alerts, and dashboards on platforms like Datadog or Splunk.

Operational Excellence

  • Production Experience: Experience supporting customer-facing production systems.
  • Problem Solving: Proven ability to logically troubleshoot complex infrastructure and application issues.
  • Reliability Concepts: Understanding of fundamental SRE concepts, such as SLOs, SLA boundaries, and blameless post-mortem cultures.
  • Continuous Integration: Familiarity with modern CI/CD pipelines and automated deployment strategies.

Security & Compliance

  • Cloud Security: Basic understanding of cloud security fundamentals, including Identity & Access Management (IAM) and secure secrets management.

Professional & Team Skills

  • Strong Collaboration: Excellent communication and collaboration skills to work effectively with globally distributed engineering teams.
  • Growth Mindset: High motivation to learn from senior engineers, accept technical feedback, and continually grow technical capabilities.
  • Execution Focus: Proven ability to deliver assigned tasks on time while maintaining work quality.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Principal Site Reliability Engineer
Principal Site Reliability Engineer

HeyGen • Tempe (AZ)

On-site
USD 140,000 - 190,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

GoGuardian • El Segundo (CA)

Hybrid
USD 180,000 - 240,000
Sr SRE Automation Engineer
Sr SRE Automation Engineer

Compunnel, Inc. • Austin (TX), Northern (KY)

On-site
USD 130,000 - 180,000
Senior Lead Site Reliability Engineer
Senior Lead Site Reliability Engineer

JPMorgan Chase & Co. • Jersey City (NJ)

On-site
USD 150,000 - 210,000
Site Reliability Engineering Manager
Site Reliability Engineering Manager

DriveWealth • Chicago (IL)

Hybrid
USD 160,000 - 230,000
Health & wellness packages
Unlimited vacation
Remote/hybrid options
+1
Senior Site Reliability Engineer, AI Agents & Automation
Senior Site Reliability Engineer, AI Agents & Automation

ServiceTitan • United States

On-site
USD 140,000 - 190,000
Flexible time off
Fully paid medical, dental, and vision
HSA/FSA programs
+7
Site Reliability Engineer
Site Reliability Engineer

TalentDome Staffing • United States

On-site
USD 140,000 - 210,000
Site Reliability Engineer
Site Reliability Engineer

Harrison Clarke • New York (NY)

On-site
USD 120,000 - 160,000
Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

Staffing Science • Arizona

On-site
USD 180,000 - 240,000
Senior SRE Engineer
Senior SRE Engineer

Compunnel, Inc. • Alpharetta (GA)

On-site
USD 140,000 - 190,000