Senior Site Reliability Engineer (CI-CD/CTAP/Delivery team)

Okta

Bengaluru

On-site

INR 4,000,000 - 7,000,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Okta is seeking an experienced Senior Site Reliability Engineer to join its Emerging Products Group in Bengaluru. The role focuses on building highly reliable, scalable cloud services with an automation-first mindset and strong emphasis on observability and operational excellence.

You will partner with software engineers, architects, and product teams to design, build, and operate production systems at scale while driving incident response, reliability metrics, and platform engineering

Qualifications

  • Experience operating large-scale production services in cloud environments.
  • Strong expertise with Kubernetes in production.
  • Proficient with Terraform and Helm.
  • Proficient in Go and/or Python.
  • Experience building automation and internal engineering platforms.
  • Experience with distributed data stores like PostgreSQL, Redis, OpenSearch, MySQL, Cassandra.

Responsibilities

  • Design, build, and operate large-scale cloud infrastructure and production services.
  • Participate in an on-call rotation for highly available customer-facing systems.
  • Lead incident response and drive post-incident reviews.
  • Define, measure, and improve SLIs, SLOs, and error budgets.
  • Collaborate with engineering teams to improve service availability, scalability, and performance.
  • Improve observability with metrics, logging, tracing, dashboards, and alerting.
  • Develop automation using Go, Python, Terraform, and related technologies.
  • Eliminate operational toil through automation and platform engineering.
  • Improve deployment safety with CI/CD and GitOps practices.
  • Collaborate on modernizing workloads and align with evolving platform capabilities.
  • Build self-service platforms and automation to boost developer velocity while maintaining reliability and security.

Skills

Kubernetes
Go
Python
Terraform
Helm
AWS
GCP
CI/CD
GitOps
Observability
Incident response
PostgreSQL
Redis
OpenSearch
MySQL
Cassandra
Networking fundamentals
Security best practices

Tools

Terraform
Helm

Job description

Secure Every Identity, from AI to Human
Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence.

Secure Every Identity, from AI to Human
Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence.
This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk.

The Engineering Opportunity

We are looking for an experienced Senior Site Reliability Engineer to join Okta's Emerging Products Group (EPG). Our mission is to build highly reliable, scalable, and secure cloud services that our customers can trust. We embrace an automation-first mindset and continuously invest in platform engineering, observability, and operational excellence to enable our engineering teams to move quickly and safely.

This role is ideal for an experienced Site Reliability Engineer who enjoys solving complex technical challenges at scale, building automation, and improving the reliability of production systems. You will serve as a key contributor within the EPG SRE organization, partnering closely with software engineers, architects, and product teams to design, build, and operate world-class cloud services.

What You'll Be Doing
Reliability & Operations
  • Design, build, and operate large-scale cloud infrastructure and production services.
  • Participate in an on-call rotation supporting highly available customer-facing systems.
  • Lead incident response efforts and drive post-incident reviews focused on systemic improvements.
  • Define, measure, and improve Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets.
  • Partner with engineering teams to improve service availability, scalability, performance, and resilience.
  • Continuously improve observability through metrics, logging, tracing, dashboards, and alerting.
  • Develop software, automation, and infrastructure using Go, Python, Terraform, and related technologies.
  • Eliminate operational toil through automation, tooling, and platform engineering.
  • Improve deployment safety and operational workflows through CI/CD and GitOps practices.
  • Collaborate on modernizing existing workloads and aligning them with evolving platform capabilities.
  • Build self-service platforms, operational guardrails, and automation that improve developer velocity while maintaining reliability and security.
What We Are Looking For
Technical Excellence
  • Strong experience operating large-scale production services in AWS and/or GCP.
  • Deep expertise with Kubernetes in production environments.
  • Experience troubleshooting Kubernetes networking, storage, scheduling, scaling, and workload lifecycle issues.
  • Extensive experience with Infrastructure as Code technologies such as Terraform and Helm.
  • Strong software engineering skills in Golang and/or Python.
  • Experience building automation and internal engineering platforms.
  • Experience operating and troubleshooting distributed data platforms such as PostgreSQL, Redis, OpenSearch, MySQL, Cassandra, or similar technologies.
  • Strong understanding of cloud networking fundamentals including DNS, load balancing, ingress, TLS, service networking, and traffic management.
  • Experience with observability platforms, monitoring strategies, and production telemetry.
  • Experience with or strong interest in AI-assisted engineering and operational automation.
Operational Excellence
  • Strong expertise operating customer-facing production systems.
  • Experience leading incident response and driving operational improvements.
  • Deep understanding of reliability engineering concepts including SLIs, SLOs, error budgets, and capacity planning.
  • Strong understanding of CI/CD pipelines, deployment strategies, and automation-first operational practices.
  • Proven ability to balance reliability, scalability, security, and engineering velocity.
The Okta Experience
  • Driving Social Impact
  • Developing Talent and Fostering Connection + Community

We are intentional about connection. Our global community, spanning over 20 offices worldwide, is united by a drive to innovate. Your journey begins with an immersive, in-person onboarding experience designed to accelerate your impact and connect you to our mission and team from day one.
Okta is an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, ancestry, marital status, age, physical or mental disability, or status as a protected veteran. We also consider for employment qualified applicants with arrest and convictions records, consistent with applicable laws.
If reasonable accommodation is needed to complete any part of the job application, interview process, or onboarding pleaseuse this Form to request an accommodation.
Notice for New York City Applicants & Employees: Okta may use Automated Employment Decision Tools (AEDT), as defined by New York City Local Law 144, that use artificial intelligence, machine learning, or other automated processes to assist in our recruitment and hiring process. In accordance with NYC Local Law 144, if you are an applicant or employee residing in New York City, pleaseclick here to view our full NYC AEDT Notice.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer (CI-CD/CTAP/Delivery team)
Senior Site Reliability Engineer (CI-CD/CTAP/Delivery team)

Engg • Bengaluru

On-site
INR 4,500,000 - 7,500,000
Staff Site Reliability Engineer - Ecosystem
Staff Site Reliability Engineer - Ecosystem

Okta • Bengaluru

On-site
INR 4,000,000 - 6,000,000
Immersive onboarding
Global community
Equal opportunity employer
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Okta • Bengaluru

On-site
INR 400,000 - 700,000
Well-being programs
Social impact initiatives
Talent development & community
Principal Site Reliability Engineer
Principal Site Reliability Engineer

Okta • Bengaluru

On-site
INR 2,000,000 - 3,000,000
Immersive onboarding experience
Equal Opportunity Employer benefits
Manager- Site Reliability Engineering
Manager- Site Reliability Engineering

Okta • Bengaluru

On-site
INR 3,000,000 - 6,000,000
Driving social impact
Talent development & community
Senior Fullstack Engineer [ Node.js Heavy + React ]
Senior Fullstack Engineer [ Node.js Heavy + React ]

Okta • Bengaluru

On-site
INR 3,500,000 - 7,000,000
The Okta Experience
Supporting Your Well-Being
Developing Talent and Fostering Connec
Senior Software Engineer
Senior Software Engineer

Okta • Bengaluru

On-site
INR 4,000,000 - 5,600,000
Well-being support
Social impact
Talent development & community
Principal Software Engineer
Principal Software Engineer

Okta • Bengaluru

On-site
INR 4,000,000 - 7,000,000
Software Backend Architect
Software Backend Architect

Engg • Bengaluru

Hybrid
INR 1,800,000 - 3,000,000
Software Engineering Manager
Software Engineering Manager

Okta • Bengaluru

On-site
INR 4,200,000 - 6,400,000
Benefits
Social impact
Talent & community