Senior Site Reliability Engineer

MangoApps

Maharashtra

On-site

INR 4,000,000 - 7,000,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

MangoApps is seeking a senior, hands-on reliability engineer to own production platform reliability, availability, security, and performance on Google Cloud. This individual-contributor role focuses on tuning infrastructure, observability, and automation, not people management.

You will drive incident response, conduct post-incident reviews, and write automation to eliminate toil. You’ll work hands-on with Docker, GKE, IAM, networking, DR planning, and collaborate with cross-functional teams to

Qualifications

  • Production-scale experience operating cloud infrastructure on Google Cloud.
  • Deep Linux systems administration, troubleshooting, and performance tuning.
  • Hands-on Docker in production.
  • Solid networking fundamentals: VPCs, routing, load balancing, DNS, VPNs, and security controls.
  • Monitoring and observability with tools such as Prometheus, Grafana, ELK/OpenSearch, Datadog, or equivalents.
  • Scripting and automation in Bash, Python, or similar.
  • Git-based workflows and CI/CD pipelines; config management with Ansible or Puppet.
  • Strong incident management and root‑cause analysis instincts, with a bias toward fixing the system, not the symptom.

Responsibilities

  • Reliability and incident response: keep the platform available and performant; define and sharpen monitoring, alerting, and observability; lead production troubleshooting and post‑incident reviews.
  • Google Cloud infrastructure and operations: design, deploy, and optimise GCP infra; drive reliability, scalability, performance, and cost; own DR and business‑continuity processes.
  • Containers and orchestration: build and operate containerised workloads on Docker with a focus on security and scaling.
  • Automation and Infrastructure as Code: provision the scripts and tooling that make operations boring and repeatable.
  • CI/CD and release engineering: maintain CI/CD pipelines and deployment automation; partner with engineering to make releases safer and easier to roll back.
  • Security and compliance: apply cloud security practices across IAM, network security, secrets management, and vulnerability remediation.

Skills

Cloud Operations
Site Reliability Engineering
Linux troubleshooting
Automation
CI/CD workflows
Observability
Incident management

Tools

Docker
Kubernetes (GKE)
Prometheus
Grafana
ELK/OpenSearch
Datadog

Job description

We're hiring a senior, hands-on engineer to own the reliability, availability, security, and performance of that platform in production, primarily on Google Cloud. This is a deep individual‑contributor role, not a management track. You'll spend your time in the systems: tuning infrastructure, building observability, automating away toil, and leading the technical response when production is on the line. You'll influence how the rest of engineering builds and operates reliable services through your work and your judgment, not through a reporting line.

If you get satisfaction from understanding a production system end to end, finding the real root cause instead of the convenient one, and making the next incident less likely, this role is built for you.

Responsibilities
  • Reliability and incident response - Keep the platform available and performant. Define and continuously sharpen monitoring, alerting, and observability. Lead production troubleshooting, drive root‑cause analysis, and run post‑incident reviews that actually change the system afterwards. Participate in the on‑call rotation for the services you own.
  • Google Cloud infrastructure and operations: Design, deploy, and optimise our GCP infrastructure - compute, storage, networking, DNS, load balancing, and security services. Drive architecture improvements for reliability, scalability, performance, and cost. Own disaster recovery and business‑continuity processes, and prove they work before you need them.
  • Containers and orchestration: Build and operate containerised workloads on Docker with a focus on security, performance, and predictable scaling across environments.
  • Automation and Infrastructure as Code Provision. Build the scripts and tooling that make operations boring and repeatable.
  • CI/CD and release engineering: Maintain CI/CD pipelines and deployment automation. Partner with engineering to make releases safer, faster, and easier to roll back.
  • Security and compliance: Apply cloud security practices across IAM, network security, secrets management, and vulnerability remediation. Keep infrastructure aligned to our internal security standards and compliance obligations.
Requirements
  • 5+ years of hands‑on Cloud Operations and Site Reliability Engineering, operating production‑scale SaaS (not pre‑production or internal‑only systems).
  • You operate Google Cloud at production scale today and can speak in specifics about GCP compute, networking, IAM, GKE, and the operational realities of running real workloads there. This is a hard requirement.
  • A second cloud (AWS or Azure) is a plus, not a substitute. We value it, but GCP depth is what the role turns on.
  • You debug Linux at the level of "why is this latency spike happening," not just "restart the service."
  • You reach for automation by reflex. Manual operational work bothers you, and you've built the tooling to remove it.
Must Have
  • Production‑scale experience operating cloud infrastructure on Google Cloud.
  • Deep Linux systems administration, troubleshooting, and performance tuning.
  • Hands‑on Docker in production.
  • Solid networking fundamentals: VPCs, routing, load balancing, DNS, VPNs, and security controls.
  • Monitoring and observability with tools such as Prometheus, Grafana, ELK/OpenSearch, Datadog, or equivalents.
  • Scripting and automation in Bash, Python, or similar.
  • Git‑based workflows and CI/CD pipelines; config management with Ansible or Puppet.
  • Strong incident management and root‑cause analysis instincts, with a bias toward fixing the system, not the symptom.
Nice To Have
  • Production experience on a second cloud (AWS or Azure).
  • IAM / SSO experience (SAML, OAuth, Okta, or similar).
  • Multi‑region or multi‑cloud operations at scale.
  • Background in cloud security, compliance, and governance practices.
What Success Looks Like
  • Sustained, high platform uptime against clear SLOs.
  • Faster incident detection and resolution, with recurring failure classes systematically driven down.
  • More automation and meaningfully less manual operational toil quarter over quarter.
  • Observability is good so that the team sees problems before customers do.
  • Engineering teams ship reliably because the operational foundation is solid.
What We're Looking For In You
  • Ownership: You take accountability for outcomes, not just tasks.
  • Problem solver: You enjoy diagnosing and resolving complex infrastructure and production challenges.
  • Continuous learner: You stay current with evolving cloud, automation, and reliability practices.
  • Collaborative: You work effectively across teams and communicate clearly, both in routine operations and in the middle of a critical incident.
  • Customer‑focused: You understand that infrastructure reliability directly shapes customer experience and business success.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer
Senior Site Reliability Engineer

MangoApps INC. • Pune District

On-site
INR 1,400,000 - 1,800,000
Site Reliability Engineer
Site Reliability Engineer

Indihire Consultants • Hyderabad

Hybrid
INR 1,500,000 - 2,800,000
Lead DevOps Engineer
Lead DevOps Engineer

WizCommerce • Bengaluru

On-site
INR 3,500,000 - 5,500,000
Senior Team Lead | Engineering, AI & Data - Engineering | Site Reliability Engineering
Senior Team Lead | Engineering, AI & Data - Engineering | Site Reliability Engineering

Deloitte & Touche GmbH Wirtschaftsprüfungsgesellschaft • Bengaluru

On-site
INR 2,000,000 - 3,000,000
Gammastack - DevOps Engineer - CI/CD Pipeline
Gammastack - DevOps Engineer - CI/CD Pipeline

GAMMASTACK • Kolkata District

On-site
INR 1,200,000 - 2,000,000
Intermediate Applications Developer
Intermediate Applications Developer

UPS • Chennai District

On-site
INR 1,500,000 - 2,000,000
Senior DevOps Engineer – GCP
Senior DevOps Engineer – GCP

MandaapX • India

On-site
INR 1,800,000 - 3,200,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

NCR Voyix • Chennai District

On-site
INR 3,000,000 - 5,400,000
Senior Site Reliability Lead
Senior Site Reliability Lead

Generac • Pune District

On-site
INR 3,000,000 - 6,500,000
Senior Infrastructure Engineer
Senior Infrastructure Engineer

Equiti Group • Bengaluru

On-site
INR 2,500,000 - 4,000,000