Lead Site Reliability Engineer

Searce Inc

Bengaluru

On-site

INR 3,500,000 - 7,000,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Searce Inc in Bengaluru seeks a Lead Site Reliability Engineer to anchor reliability and scalability for client cloud environments, primarily on Google Cloud Platform with exposure to AWS or Azure. You will drive SRE practices, mentor engineers, and partner with DevOps and platform squads to engineer resilient systems that meet SLA commitments.

You will own SLOs/SLIs, incident response, and chaos engineering, while advancing auto‑healing, multi‑cloud resilience, and GitOps‑driven deployments.

Qualifications

  • 6–8 years of experience in SRE, DevOps, or cloud infrastructure roles.
  • GCP as primary cloud—GKE, VPC, IAM, Monitoring, Pub/Sub, BigQuery.
  • Minimum 2 cloud platforms certifications or equivalent production experience across 2 clouds.
  • Infrastructure‑as‑Code expertise: Terraform, Pulumi, or Cloud Deployment Manager.
  • SRE practice: SLO/SLI definition, incident management, toil reduction.
  • Scripting: Python, Go, Bash for runbooks and tooling.

Responsibilities

  • Own SLOs, SLIs, and error budgets; drive reliability roadmaps.
  • Lead incident response, post‑mortems, and blameless RCA practices; reduce toil via automation.
  • Design auto‑healing, self‑remediation, and chaos engineering frameworks.
  • Ensure high availability, DR, and resilience across multi‑clouds.
  • Architect and operate production‑grade GCP environments (GKE, Cloud Run, BigQuery, Pub/Sub).
  • Extend multi‑cloud coverage on AWS or Azure as a secondary platform.
  • Manage IaC using Terraform/ Pulumi/ Deployment Manager; enforce GitOps.

Skills

SRE practices
GCP expertise
Multi-cloud
Cloud certifications
Terraform / IaC
Scripting (Python/Go/Bash)
Observability tooling

Tools

Terraform
Pulumi
Cloud Deployment Manager

Job description

Searce is an AI outcome engineering company that helps businesses transform through deep technology expertise. With 3,500+ engineers spread across 12 countries, we partner with enterprises to design, build, and operate cloud-native, AI-first systems that generate measurable business outcomes. Searce is a Google Cloud Global MSP, AWS Advanced Consulting Partner, Databricks Partner, and Anthropic Claude Partner. We are PCI DSS and ISO 27001 certified, ensuring enterprise-grade security and compliance.

Role overview

As a Lead SRE within the CSRE Delivery team, you will anchor reliability, scalability, and operational excellence for our client cloud environments — primarily on GCP, with hands‑on exposure to at least one additional cloud platform (AWS or Azure). You will drive SRE practices, mentor engineers, and partner closely with DevOps and platform squads to engineer resilient systems that exceed SLA commitments.

Key responsibilities
SRE & Reliability Engineering
  • Own SLOs, SLIs, and error budget policies; drive reliability roadmaps for client workloads.
  • Lead incident response, post-mortems, and blameless RCA practices; eliminate toil through automation.
  • Design and implement auto-healing, self-remediation, and chaos engineering frameworks.
  • Ensure high availability, disaster recovery, and resilience across multi-cloud environments.
  • Architect and operate production‑grade GCP environments (GKE, Cloud Run, BigQuery, Pub/Sub, etc.).
  • Extend multi‑cloud coverage on AWS or Azure as a secondary platform (minimum 2 cloud platforms required).
  • Manage infrastructure-as-code using Terraform / Deployment Manager; enforce GitOps workflows.
DevOps & CI/CD
  • Own and optimise CI/CD pipelines (Cloud Build, Jenkins, ArgoCD, GitHub Actions, or equivalent).
  • Champion container‑native delivery on Kubernetes / GKE; define deployment strategies (blue‑green, canary).
  • Integrate security and compliance gates (SAST, DAST, policy‑as‑code) within pipeline workflows.
Observability & Performance
  • Build end-to-end observability stacks: Cloud Monitoring, Prometheus, Grafana, Datadog, or equivalent.
  • Define alerting strategies, dashboards, and runbooks; conduct capacity planning and performance tuning.
  • Drive AIOps and intelligent alerting adoption to reduce MTTR.
  • Mentor and guide a team of SRE/DevOps engineers; conduct code and design reviews.
  • Collaborate with Delivery Managers, architects, and client stakeholders on SRE roadmaps.
  • Contribute to CSRE capability building, tooling standards, and internal knowledge assets.
Skills & requirements
Mandatory
  • 6 – 8 years of total experience in SRE, DevOps, or cloud infrastructure roles.
  • Strong hands‑on expertise in Google Cloud Platform (GCP) as primary cloud — GKE, VPC, IAM, Cloud Monitoring, Pub/Sub, BigQuery, Cloud SQL.
  • Proficiency in at least one secondary cloud: AWS (EKS, EC2, RDS, CloudWatch) or Azure (AKS, Azure Monitor, ARM).
  • Minimum 2 cloud platform certifications or equivalent production experience across 2 clouds.
  • Infrastructure‑as‑Code expertise: Terraform, Pulumi, or Cloud Deployment Manager.
  • Demonstrated SRE practice: SLO/SLI definition, error budgets, incident management, toil reduction.
  • Scripting / automation: Python, Go, Bash — for runbooks, tooling, and ops automation.
  • Observability tooling: Prometheus, Grafana, Cloud Monitoring, Datadog, or PagerDuty.
Good to have
  • Experience with service mesh (Istio, Anthos Service Mesh).
  • Familiarity with FinOps practices and cloud cost optimisation.
  • Exposure to AIOps platforms or ML‑based anomaly detection tools.
  • Experience in a managed services or cloud consultancy delivery environment.
What we offer
  • Work on cutting‑edge, large‑scale GCP‑first client environments across industries.
  • Access to Searce's AI outcome engineering learning ecosystem and certifications.
  • Cross‑cloud exposure — GCP, AWS, and Azure — within a single delivery team.
  • Transparent growth framework with clear paths to Senior Manager / AVP tracks.
  • Collaborative, engineer‑led culture with direct client impact and ownership.
About the CSRE team

The Cloud Site Reliability Engineering (CSRE) Delivery team at Searce is responsible for designing, running, and evolving mission‑critical cloud operations for enterprise clients globally. The team embeds SRE discipline into every delivery — from onboarding to steady‑state ops — ensuring clients benefit from proactive reliability engineering, not just reactive support.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Lead Site Reliability Engineer (SRE)
Lead Site Reliability Engineer (SRE)

Skillventory • Kamrup Metropolitan

On-site
INR 3,500,000 - 7,000,000
Senior Site Reliability Engineer Cloud Ops Hyderabad
Senior Site Reliability Engineer Cloud Ops Hyderabad

Seismic • Hyderabad

On-site
INR 2,500,000 - 4,000,000
Sr. Site Reliability Engineer-GCP,Terraform,Networking,Linux ,Python
Sr. Site Reliability Engineer-GCP,Terraform,Networking,Linux ,Python

Optum • Hyderabad

Hybrid
INR 1,800,000 - 2,800,000
Lead SRE
Lead SRE

Cvent, Inc. • Gurugram District

On-site
INR 4,000,000 - 8,000,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Seismic • Hyderabad

On-site
INR 4,000,000 - 7,500,000
Site Reliability Engineer- GCP
Site Reliability Engineer- GCP

Aziro • Hyderabad

On-site
INR 1,400,000 - 2,100,000
Senior Site Reliability Engineer (SRE) – GCP
Senior Site Reliability Engineer (SRE) – GCP

Opstree Global • Bengaluru

On-site
INR 3,600,000 - 6,000,000
Lead Engineer – Site Reliability Engineering
Lead Engineer – Site Reliability Engineering

CBTS • Chennai District

On-site
INR 5,000,000 - 7,500,000
Lead SRE
Lead SRE

Cvent, Inc. • India

On-site
INR 2,500,000 - 4,500,000
Lead Site Reliability Engineer
Lead Site Reliability Engineer

Sierra Ventures • Bengaluru

On-site
INR 3,500,000 - 5,500,000