Site Reliability Engineer

Mumba Technologies, Inc.

Gurugram District

Hybrid

INR 2,500,000 - 4,000,000

Full time

33 hours ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Mumba Technologies, Inc. in Gurgaon is seeking an experienced Site Reliability Engineer to build and operate highly reliable, scalable, secure, and observable systems in a hybrid work model with 3 days in the office.

You will lead incident response, design comprehensive observability, and implement AI-assisted reliability practices while mentoring junior SREs. The role spans AWS, Kubernetes, IaC, security, and performance optimization.

Qualifications

  • Lead incident response, RCAs and track actions to closure.
  • Design and own observability stack across metrics, logs, traces and events.
  • Drive toil reduction by eliminating repetitive ops work.

Responsibilities

  • Lead incident response, conduct RCAs and ensure action items are tracked to closure.
  • Build and maintain runbooks, playbooks and escalation frameworks for proactive and reactive response.
  • Drive toil reduction by identifying repetitive operational work and engineering it away.
  • Design and own the full observability stack — metrics, logs, traces and events using Prometheus, Grafana, OpenTelemetry, Datadog, Dynatrace.
  • Build intelligent alerting that reduces noise, surfaces actionable signals.
  • Implement distributed tracing and dependency mapping for end-to-end visibility across microservices.
  • Leverage AI/ML for anomaly detection, predictive alerting and root-cause analysis.
  • Own Kubernetes cluster operations — autoscaling, networking, upgrades, cost optimization.
  • Mentor SREs and drive engineering standards and best practices.

Skills

Incident response
Observability design
Kubernetes operations
AWS
AI-driven SRE practices
Mentoring
SRE leadership
Toil reduction
Cost optimisation
CI/CD

Tools

Prometheus
Grafana
OpenTelemetry
Datadog
Dynatrace
Cloudflare CDN
Terraform
Pulumi
Kubernetes
AWS

Job description

Location: Gurgaon (Hybrid 3-Days in the office)
About the Role:

We are looking for an experienced Site Reliability Engineer (SRE) to help build and operate highly reliable, scalable, secure, and observable systems. The role involves hands-on work across AWS, Kubernetes, Infrastructure as Code, observability, automation, security, and AI-driven SRE practices.

Reliability & Availability
  • Lead incident response, conduct RCAs and ensure action items are tracked to closure
  • Build and maintain runbooks, playbooks and escalation frameworks for proactive and reactive response
  • Drive toil reduction by identifying repetitive operational work and engineering it away
Observability
  • Design and own the full observability stack — metrics, logs, traces and events — using tools like Prometheus, Grafana, OpenTelemetry, Datadog, Dynatrace or similar
  • Build intelligent alerting that reduces noise, eliminates alert fatigue and surfaces actionable signals
  • Implement distributed tracing and dependency mapping to provide end-to-end visibility across microservices
  • Drive adoption of continuous profiling and real user monitoring (RUM) for proactive performance management
AI Adoption in SRE
  • Leverage AIOps platforms to enable anomaly detection, predictive alerting and automated root cause analysis
  • Implement AI-assisted incident triage — using LLM-powered tools to summarise incidents, suggest fixes and accelerate MTTR
  • Build and maintain ML-powered capacity forecasting models to optimise infrastructure spend and prevent resource saturation
Security & Vulnerability Management
  • Embed security-as-reliability principles — treating security incidents with the same urgency as availability incidents
  • Own issue remediations across infrastructure (OS, containers, dependencies)
  • Integrate SAST, DAST and SCA tools into CI/CD pipelines to shift security left
  • Design, build and maintain cloud-native infrastructure on AWS using Infrastructure as Code (Terraform, Pulumi) & drive rightsizing, reserved capacity planning and cost anomaly detection
  • Own Kubernetes cluster operations — autoscaling, resource management, networking and upgrade strategy
Leadership & Culture
  • Mentor and guide junior and mid-level SREs — conducting technical reviews and pair debugging sessions
  • Define and evolve SRE team standards, best practices and engineering principles
  • Collaborate closely with product, development and security teams as an embedded reliability partner
  • Contribute to on-call rotation and drive continuous improvement of on-call experience
  • Represent SRE in architecture reviews, sprint planning and cross-functional forums
The Competitive Edge
AEM Administration
  • Own end-to-end reliability and availability of AEM environments — Author, Publish, Dispatcher and AEM as a Cloud Service (AEMaaCS) — across dev, staging and production
  • Monitor and manage AEM instance health, optimise Dispatchers, Manage DAM, OSGi Configurations, replication queues.
Exposure to CDN
  • Experience with Cloudflare - CDN, Workers.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

SourcingXPress • Maharashtra

On-site
INR 700,000 - 1,800,000
Site Reliability Engineer
Site Reliability Engineer

Infitoo Systems • Dadri

On-site
INR 2,400,000 - 3,600,000
Site Reliability Engineer
Site Reliability Engineer

SourcingXPress • Mumbai

On-site
INR 800,000 - 1,200,000
Senior Site Reliability Engineer - Cloud Infrastructure
Senior Site Reliability Engineer - Cloud Infrastructure

WITS Innovation Lab • Chandigarh

On-site
INR 1,800,000 - 3,000,000
Site Reliability Engineer
Site Reliability Engineer

Arch Systems • Hyderabad

On-site
INR 2,800,000 - 4,200,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

VMC Soft Technologies, Inc • Hyderabad

Hybrid
INR 1,500,000 - 2,000,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

AcquireX • Pune District

On-site
INR 1,200,000 - 1,800,000
Health insurance
Flexible working hours
Training opportunities
Senior Web Operations Engineer (SRE)
Senior Web Operations Engineer (SRE)

Keysight Technologies SAles Spain SL. • Gurugram District

On-site
INR 4,000,000 - 7,000,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Infosys • Bengaluru

On-site
INR 900,000 - 1,500,000
Resilience and Reliability Engineer
Resilience and Reliability Engineer

EY • Pune District, Gurugram District, Bengaluru

Hybrid
INR 1,800,000 - 2,800,000