Site Reliability Engineer

Razorpay

Bengaluru

On-site

INR 4,000,000 - 7,000,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Razorpay in Bengaluru seeks a founding infrastructure SRE to own reliability of the platform across compute, orchestration, networking, data stores, CI/CD and observability.

As a one-person foundation, you will standardize health checks, dashboards, and failure-domain design, lead incident response, and drive scalable automation for product teams.

You will collaborate with engineering to embed reliability, plan capacity for peak events, and shape the on-call model for sustainable coverage.

Qualifications

  • 12+ years owning production infrastructure at scale.
  • Deep Kubernetes and container orchestration in production.
  • Strong Linux internals and networking expertise.
  • Experience with at least one major cloud and IaC (Terraform).
  • Track record building platform tooling adopted by multiple teams.
  • Genuine on-call and incident command experience on infrastructure used by many teams.

Responsibilities

  • Own reliability of core infrastructure: Kubernetes clusters, networking and ingress, DNS, databases and caches, message queues, and the CI/CD platform.
  • Build the paved road: golden paths for deployment, standard health checks, default dashboards and alerts, and production readiness templates that product teams adopt because they are the easiest option.
  • Design and test failure domains: multi-AZ and multi-region strategy, failover automation, capacity planning, and disaster recovery drills that are actually run, not just documented.
  • Own infrastructure-level rollouts and rollbacks: cluster upgrades, database migrations, and platform changes executed without customer-visible impact.
  • Build and tune the observability platform so that every team can see its own SLIs without asking for help.
  • Carry the pager for platform services, lead infrastructure incidents, and drive systemic fixes across teams.
  • Establish a shared on-call model with existing platform engineers so that infrastructure coverage does not depend on one person.
  • Drive cost-aware capacity management: headroom for peak events without permanent overprovisioning.

Skills

Kubernetes
Linux internals
Cloud (AWS)
Terraform
Incident management
Platform tooling
Capacity planning
Multi-tenant infra
Networking depth

Job description

You will be the founding infrastructure SRE at Razorpay, responsible for the reliability of the platform every product team builds on: compute, orchestration, networking, data stores, CI/CD, and observability. Payment platform SREs make individual flows reliable; you make the ground they stand on reliable. Because you are one person covering a wide surface, your primary weapon is leverage: paved roads, standards, and automation that let service teams own their own reliability.

What you will do
  • Own the reliability of core infrastructure: Kubernetes clusters, networking and ingress, DNS, databases and caches, message queues, and the CI/CD platform.
  • Build the paved road: golden paths for deployment, standard health checks, default dashboards and alerts, and production readiness templates that product teams adopt because they are the easiest option.
  • Design and test failure domains: multi-AZ and multi-region strategy, failover automation, capacity planning, and disaster recovery drills that are actually run, not just documented.
  • Own infrastructure-level rollouts and rollbacks: cluster upgrades, database migrations, and platform changes executed without customer-visible impact.
  • Build and tune the observability platform so that every team can see its own SLIs without asking for help.
  • Carry the pager for platform services, lead infrastructure incidents, and drive systemic fixes across teams.
  • Establish a shared on-call model with existing platform engineers so that infrastructure coverage does not depend on one person.
  • Drive cost-aware capacity management: headroom for peak events (festival sales, month-end settlement spikes) without permanent overprovisioning.
What we are looking for
  • 12+ years of experience, with substantial time owning production infrastructure at a company with real scale (thousands of hosts or containers, or traffic with hard peaks).
  • Deep Kubernetes and container orchestration experience in production, including upgrades, capacity, and multi-tenant reliability.
  • Strong Linux internals and networking depth: you can debug from a TCP dump, a flame graph, or kernel-level metrics, not just from a dashboard.
  • Production experience with at least one major cloud (AWS preferred) and infrastructure as code (Terraform or similar).
  • Experience operating stateful systems under load: relational databases, caches, and queues, including failover and data migration without downtime.
  • A track record of building platform tooling or standards adopted by other teams, because a team of one scales only through software and influence.
  • Genuine on-call and incident command experience on infrastructure that many teams depended on.
  • The temperament for a founding role: comfortable with ambiguity, able to prioritize ruthlessly, and able to say no with data.
Nice to have
  • Experience running infrastructure for payments, banking, or another regulated, correctness-critical domain.
  • Chaos engineering or failure injection experience at the infrastructure layer.
  • Experience planning for extreme traffic events (large sale days, live sports streaming, ticket sales).
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

Razorpay Software Pvt Ltd • Bengaluru

On-site
INR 3,500,000 - 7,500,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Impronics Technologies • Gurugram District

On-site
INR 1,500,000 - 2,500,000
Site Reliability Engineer
Site Reliability Engineer

Recro • Bengaluru

On-site
INR 3,500,000 - 6,500,000
Site Reliability Engineer
Site Reliability Engineer

SourcingXPress • Mumbai

On-site
INR 800,000 - 1,200,000
Site Reliability Engineer
Site Reliability Engineer

Spot Your Leaders & Consulting • Pune District

On-site
INR 2,500,000 - 4,000,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Bahwan CyberTek • Hyderabad

On-site
INR 1,200,000 - 1,800,000
Lead Site Reliability Engineer
Lead Site Reliability Engineer

Sierra Ventures • Bengaluru

On-site
INR 3,500,000 - 5,500,000
Site Reliability Engineer
Site Reliability Engineer

Smart Ims • Bengaluru

On-site
INR 1,200,000 - 2,000,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Embarkgcc Services • Bengaluru

On-site
INR 1,200,000 - 1,800,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Falabella India • Bengaluru

On-site
INR 4,000,000 - 7,000,000