Site Reliability Engineer

Insight Global

Bengaluru

On-site

INR 2,500,000 - 5,000,000

Full time

6 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Insight Global is seeking a Kubernetes Site Reliability Engineer to keep Ecolab’s container platform reliable and secure across Azure and AWS.

Applying software engineering to operations, you’ll set service-level objectives, lead incident response and on-call, automate toil, and help migrate Azure-native services toward cloud-agnostic patterns—codifying everything as Infrastructure as Code, building safe delivery pipelines, and hardening to CIS Benchmarks.

Qualifications

  • 5+ years in SRE, infrastructure, or platform engineering with hands-on Kubernetes.
  • Deep AWS and Azure (services, networking, IAM) and strong Kubernetes / EKS.
  • Reliability fundamentals - SLIs/SLOs/error budgets, incident response, on-call.
  • Observability tooling (Elastic plus Prometheus/Grafana/Datadog).
  • Terraform multi-cloud IaC; scripting in Python, Bash, and YAML.
  • Cloud migration - moving Azure-native services (Functions, Cosmos DB) toward AWS / cloud-agnostic.
  • Policy and access control - OPA/Gatekeeper, AWS IAM, Azure RBAC.
  • Kubernetes troubleshooting - capacity planning, performance tuning.

Responsibilities

  • Reliability & SLOs: Define and maintain SLIs, SLOs, and error budgets, and use them to drive priorities.
  • Incident & on-call: Lead incident response on-call - detect, triage, resolve - then run blameless post-mortems.
  • Observability: Build and tune monitoring, logging, tracing, and alerting (Elastic, Prometheus, Grafana, Datadog).
  • Cloud migration: Help migrate Azure-native services (Functions, Cosmos DB) toward AWS / cloud-agnostic patterns with failover, replication, and latency tuning.
  • Automation & performance: Engineer toil away (remediation, scaling, patching, backups) and run capacity planning, performance tuning, and load/scalability testing.
  • Safe delivery: Build and safeguard CI/CD and progressive delivery (Azure DevOps, GitHub Actions, ArgoCD, Argo Rollouts) with automated rollbacks.
  • Security & policy: Harden to CIS Benchmarks and apply policy and access controls - OPA/Gatekeeper, AWS IAM, Azure RBAC - remediating Mythos vulnerabilities with security.

Skills

Kubernetes
AWS
Azure
SRE
Observability
Terraform
Python
Bash
CI/CD
On-call

Tools

Elastic
Prometheus
Grafana
Datadog
ArgoCD
Argo Rollouts

Job description

  • • 5+ years in SRE, infrastructure, or platform engineering with hands-on production Kubernetes.
  • • Deep AWS and Azure (services, networking, IAM) and strong Kubernetes / EKS.
  • • Reliability fundamentals - SLIs/SLOs/error budgets, incident response, on-call.
  • • Observability tooling (Elastic plus Prometheus/Grafana/Datadog).
  • • Terraform multi-cloud IaC; scripting in Python, Bash, and YAML.
  • • Cloud migration - moving Azure-native services (Functions, Cosmos DB) toward AWS / cloud-agnostic.
  • • Policy and access control - OPA/Gatekeeper, AWS IAM, Azure RBAC.
  • • Kubernetes troubleshooting - capacity planning, performance tuning, load/scalability testing.
Nice to Have Skills & Experience
  • • Managed Kubernetes (AKS, EKS, GKE) and service mesh
  • • Argo Rollouts / Flagger progressive delivery
  • • Chaos engineering and hybrid-architecture DR patterns
  • • Kyverno and supply-chain security
Job Description

Insight Global is seeking a Kubernetes Site Reliability Engineer to keep Ecolab’s new container platform reliable, fast, and secure across Azure and AWS. Applying software engineering to operations, you’ll set service-level objectives, lead incident response and on-call, automate toil, and help migrate Azure-native services toward cloud-agnostic patterns - codifying everything as Infrastructure as Code, building safe delivery pipelines, and hardening to CIS Benchmarks. Success is measured in uptime, fast recovery, and toil removed.

Key Responsibilities:
  • • Reliability & SLOs: Define and maintain SLIs, SLOs, and error budgets, and use them to drive priorities.
  • • Incident & on-call: Lead incident response on-call - detect, triage, resolve - then run blameless post-mortems.
  • • Observability: Build and tune monitoring, logging, tracing, and alerting (Elastic, Prometheus, Grafana, Datadog).
  • • Cloud migration: Help migrate Azure-native services (Functions, Cosmos DB) toward AWS / cloud-agnostic patterns with failover, replication, and latency tuning.
  • • Automation & performance: Engineer toil away (remediation, scaling, patching, backups) and run capacity planning, performance tuning, and load/scalability testing.
  • • Safe delivery: Build and safeguard CI/CD and progressive delivery (Azure DevOps, GitHub Actions, ArgoCD, Argo Rollouts) with automated rollbacks.
  • • Security & policy: Harden to CIS Benchmarks and apply policy and access controls - OPA/Gatekeeper, AWS IAM, Azure RBAC - remediating Mythos vulnerabilities with security.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

Innodata Inc. • India

On-site
INR 2,400,000 - 4,000,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Falabella India • Bengaluru

On-site
INR 4,000,000 - 7,000,000
Staff Site Reliability Engineer
Staff Site Reliability Engineer

PowerToFly • Gurgaon

On-site
INR 1,800,000 - 3,000,000
Staff Site Reliability Engineer
Staff Site Reliability Engineer

Okta • Bengaluru

On-site
INR 1,500,000 - 2,500,000
Staff Site Reliability Engineer
Staff Site Reliability Engineer

Stryker Group • Gurugram District

On-site
INR 1,200,000 - 1,800,000
Senior Site Reliability Lead
Senior Site Reliability Lead

Generac • Pune District

On-site
INR 3,000,000 - 6,500,000
SRE Engineer
SRE Engineer

Prodapt Solutions Private Limited • Chennai District

On-site
INR 1,800,000 - 3,000,000
Site Reliability Engineer
Site Reliability Engineer

DeepIQ • Hyderabad

On-site
INR 1,200,000 - 2,100,000
Site Reliability Engineer
Site Reliability Engineer

MishiPay • Bengaluru

On-site
INR 2,500,000 - 3,800,000
Lead Site Reliability Engineer
Lead Site Reliability Engineer

Sierra Ventures • Bengaluru

On-site
INR 3,500,000 - 5,500,000