- • 5+ years in SRE, infrastructure, or platform engineering with hands-on production Kubernetes.
- • Deep AWS and Azure (services, networking, IAM) and strong Kubernetes / EKS.
- • Reliability fundamentals - SLIs/SLOs/error budgets, incident response, on-call.
- • Observability tooling (Elastic plus Prometheus/Grafana/Datadog).
- • Terraform multi-cloud IaC; scripting in Python, Bash, and YAML.
- • Cloud migration - moving Azure-native services (Functions, Cosmos DB) toward AWS / cloud-agnostic.
- • Policy and access control - OPA/Gatekeeper, AWS IAM, Azure RBAC.
- • Kubernetes troubleshooting - capacity planning, performance tuning, load/scalability testing.
Nice to Have Skills & Experience
- • Managed Kubernetes (AKS, EKS, GKE) and service mesh
- • Argo Rollouts / Flagger progressive delivery
- • Chaos engineering and hybrid-architecture DR patterns
- • Kyverno and supply-chain security
Job Description
Insight Global is seeking a Kubernetes Site Reliability Engineer to keep Ecolab’s new container platform reliable, fast, and secure across Azure and AWS. Applying software engineering to operations, you’ll set service-level objectives, lead incident response and on-call, automate toil, and help migrate Azure-native services toward cloud-agnostic patterns - codifying everything as Infrastructure as Code, building safe delivery pipelines, and hardening to CIS Benchmarks. Success is measured in uptime, fast recovery, and toil removed.
Key Responsibilities:
- • Reliability & SLOs: Define and maintain SLIs, SLOs, and error budgets, and use them to drive priorities.
- • Incident & on-call: Lead incident response on-call - detect, triage, resolve - then run blameless post-mortems.
- • Observability: Build and tune monitoring, logging, tracing, and alerting (Elastic, Prometheus, Grafana, Datadog).
- • Cloud migration: Help migrate Azure-native services (Functions, Cosmos DB) toward AWS / cloud-agnostic patterns with failover, replication, and latency tuning.
- • Automation & performance: Engineer toil away (remediation, scaling, patching, backups) and run capacity planning, performance tuning, and load/scalability testing.
- • Safe delivery: Build and safeguard CI/CD and progressive delivery (Azure DevOps, GitHub Actions, ArgoCD, Argo Rollouts) with automated rollbacks.
- • Security & policy: Harden to CIS Benchmarks and apply policy and access controls - OPA/Gatekeeper, AWS IAM, Azure RBAC - remediating Mythos vulnerabilities with security.