Senior Site Reliability Engineer

RapidAI

Bengaluru

On-site

INR 4,000,000 - 7,000,000

Full time

6 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

RapidAI is seeking a senior Site Reliability Engineer (SRE) to own the availability and performance of our production EKS clusters and observability stack. You will define SLOs/SLIs, drive post-mortems, and implement infrastructure-as-code with Terraform, Helm, and GitOps.

You will collaborate with engineering to bake reliability in early, perform capacity planning and chaos testing, and optimize costs and autoscaling across AWS workloads. Strong scripting in Go/Python/Bash is expected.

Qualifications

  • 10+ years in SRE/DevOps or infrastructure roles.
  • Deep AWS expertise: EKS/EC2/VPC/IAM/RDS/S3/CloudWatch.
  • Production Kubernetes experience at scale.
  • Hands-on Open Telemetry instrumentation and pipelines.
  • Strong Linux, networking, and distributed systems fundamentals.
  • Experience with observability platforms (Prometheus/Grafana/Jaeger).
  • Comfort with Go, Python, or Bash automation.

Responsibilities

  • Own availability, performance, and incident response for Rapid's production EKS clusters.
  • Design and operate the full observability stack (metrics, logs, traces).
  • Define and track SLOs/SLIs and lead blameless post-mortems.
  • Build infrastructure-as-code using Terraform, Helm, and GitOps patterns.
  • Collaborate with engineering to bake reliability in (capacity planning, load testing).
  • Tune autoscaling, networking, and cost efficiency across AWS workloads.

Skills

AWS expertise
Kubernetes
Open Telemetry
Linux fundamentals
Go/Python scripting
Chaos engineering
Cost optimization

Tools

Terraform
Helm
GitOps
Prometheus
Grafana
Jaeger
CloudWatch
K8s multi-cluster

Job description

RapidAI is the trusted leader in deep clinical AI, helping hospitals deliver faster, more informed care through intelligent imaging and integrated workflows. The Rapid Enterprise™ Platform supports disease states across the care spectrum, but it’s our clinical depth that drives the most meaningful impact — improving decision-making, patient outcomes, and health-system performance. Used by more than 2,500 hospitals in over 100 countries and backed by 700+ clinical studies, including research that helped expand national stroke-treatment guidelines, RapidAI is the most clinically validated AI platform in healthcare.

What You Do:
  • Own the availability, performance, and incident response for Rapid's production EKS clusters
  • Design and operate the full observability stack — metrics, logs, traces — with
    Open Telemetry as the foundation
  • Define and track SLOs/SLIs/error budgets; lead post-mortems and drive blameless culture
  • Build and maintain infrastructure-as-code using Terraform, Helm, and GitOps patterns
  • Partner with engineering to bake reliability in early — capacity planning, load testing, chaos engineering
  • Tune autoscaling, networking, and cost efficiency across AWS workloads
  • On-call rotation with the expectation you'll also fix the underlying cause, not just the alert
What We Looking For:
  • 10+ years in SRE, DevOps, or infrastructure engineering roles
  • Deep AWS expertise — EKS, EC2, VPC, IAM, RDS, S3, CloudWatch, and the
    surrounding ecosystem
  • Production Kubernetes experience at scale: multi-cluster, multi-tenant, real traffic
  • Hands-on Open Telemetry instrumentation and pipeline ownership (collectors, exporters, backends)
  • Strong foundation in Linux, networking, and distributed systems fundamentals
  • Experience with observability platforms (Prometheus, Grafana, Jaeger, or equivalents)
    Comfortable writing automation in Go, Python, or Bash — you reach for code when theGUI runs out
  • Startup mindset: you make decisions with incomplete information and iterate quickly

RapidAI is committed to creating an inclusive and diverse workplace. We provide equal employment opportunities to all employees and applicants and prohibit discrimination and harassment of any type in regard to race, color, religion, age, sex, national origin, disability status, genetics, protected veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by federal, state or local laws.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff Software Engineer Golang
Staff Software Engineer Golang

RapidAI • Bengaluru

On-site
INR 4,000,000 - 7,000,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

SourcingXPress • Hyderabad

On-site
INR 3,000,000 - 5,000,000
Staff Software Engineer, Golang
Staff Software Engineer, Golang

RapidAI • Bengaluru

On-site
INR 3,500,000 - 5,500,000
Senior SRE
Senior SRE

CloudRaft • India

On-site
INR 2,500,000 - 4,500,000
Competitive salary
Premium health insurance & wellness
AI stack & GPU infrastructure
+2
Principal Site Reliability Engineer
Principal Site Reliability Engineer

Namely • India

On-site
INR 1,500,000 - 2,500,000
AI SRE/ AI Site Reliability Engineer
AI SRE/ AI Site Reliability Engineer

Tata Consultancy Services • Bengaluru

On-site
INR 1,800,000 - 2,800,000
Forward Deployment Engineer (SRE)
Forward Deployment Engineer (SRE)

PwC Acceleration Centers • Hyderabad

On-site
INR 2,500,000 - 4,000,000
Staff Site Reliability Engineer
Staff Site Reliability Engineer

Okta • Bengaluru

On-site
INR 1,500,000 - 2,500,000
Site Reliability Engineer
Site Reliability Engineer

United States Digital Space LLC • Karnataka

On-site
INR 900,000 - 1,200,000
Significant equity in a venture-backed company
Opportunity to work with modern tech stack
Senior Manager - Site Reliability Engineer|NR-2026-0246
Senior Manager - Site Reliability Engineer|NR-2026-0246

Media.net • Bengaluru

On-site
INR 6,000,000 - 8,000,000