Lead Site Reliability Engineer

Zeta Global

Bengaluru

On-site

INR 1,200,000 - 2,400,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Zeta Global is seeking an experienced SRE to drive reliability across cloud and on-prem environments in Bengaluru. The role focuses on building resilient systems, managing SLOs/SLIs, and leading incident response with blameless postmortems.

You will implement observability using OpenTelemetry, work with AWS and Kubernetes stacks, and advance automation via IaC tools like Terraform or Pulumi. Strong coding in Python/Go and a passion for performance are essential.

Qualifications

  • 3–5 years of experience as an SRE in cloud-based and on-prem environments.
  • Deep understanding of Linux systems, networking, and sysadmin tasks.
  • Experience with AWS and Kubernetes orchestration and container tooling.
  • Hands-on experience with observability tools: Honeycomb, Grafana, Prometheus, Thanos, ELK/Loki.
  • Strong programming in Python or Go for production code.
  • Solid shell scripting (bash or similar).
  • Experience with OpenTelemetry or distributed tracing integration.
  • Experience with Chaos Engineering tools and practices.

Responsibilities

  • Implement and manage SLOs, SLIs, and error budgets to guide reliability efforts.
  • Develop systems resilient to failures, targeting 99.9%+ uptime for critical services.
  • Lead incident response and postmortems with root-cause analysis for continuous improvement.
  • Automate incident detection and response via runbooks or workflows.
  • Write code to support reliability or efficiency needs as required.
  • Design and implement full observability across systems using OpenTelemetry for tracing, metrics, and logging.
  • Plan capacity and performance testing to ensure scalable systems.
  • Collaborate with Dev and Ops to build reliable, scalable services.
  • Adopt best practices in infra design, deployment, and maintenance using AWS, Kubernetes, EKS, Fargate.
  • Champion infrastructure as code with Terraform or Pulumi for provisioning and scaling.
  • Participate in chaos engineering initiatives and on-call rotations.
  • Drive advanced alerting and anomaly detection on metrics.

Skills

Linux systems
Networking
AWS
Kubernetes
Python
Bash
OpenTelemetry
Prometheus
CI/CD
Incident management
On-call
Chaos engineering

Tools

Grafana
Prometheus
ELK
Loki
Thanos
Terraform
Pulumi

Job description

Responsibilities

Implement and manage Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error budgets to drive reliability efforts.

Develop systems that are resilient to failures and ensure 99.9%+ uptime for critical services.

Lead incident response and post‑incident reviews (blameless postmortems), ensuring robust root cause analysis and continuous improvement of systems.

Automate incident detection and response using automated runbooks or predefined workflows.

Write software as needed to support reliability or efficiency needs.

Design and implement full observability across systems using modern tools like Open Telemetry for tracing, metrics, and logging.

Use capacity planning, forecasting, and performance testing to ensure that the systems scale effectively as the user base and load grow.

Collaborate with development and operations teams on building reliable, scalable, and high‑performance services.

Ensure best practices are followed across infrastructure design, deployment, and maintenance using tools like AWS, Kubernetes, EKS, Fargate, etc.

Champion Infrastructure as Code (IaC) to provision, manage, and scale infrastructure using tools like Terraform, Pulumi, or similar.

Get involved in chaos engineering initiatives.

Participate on our on‑call rotation.

Drive advanced alerting and anomaly detection applied to metrics.

Experience & Qualifications
  • 3‑5 years of experience as an SRE, working in cloud‑based environments and on‑prem environments.
  • Deep understanding of Linux systems, networking, and systems administration.
  • Experience with cloud platforms like AWS, with a strong understanding of Kubernetes and container orchestration tools.
  • Hands‑on experience with observability tools such as Honeycomb, Grafana, Prometheus, Thanos, ELK (Elastic Stack), or Loki.
  • Strong skills in at least one programming language (Python, Go) to write production‑level code.
  • Strong skills in shell scripting using bash or similar.
  • Experience with OpenTelemetry or other distributed tracing systems, including tracing, metrics, and logs integration.
  • Experience with Chaos Engineering methodologies and tools (Chaos Mesh, Chaos Monkey, AWS Fault Injection Simulator, etc.).
Skills & Knowledge
  • Reliability‑focused mindset with the ability to balance fast product iterations and system stability.
  • Solid understanding of SLOs, SLIs, and error budgets.
  • Hands‑on knowledge of CI/CD pipelines and infrastructure automation.
  • Proven expertise in incident management, postmortems, and root cause analysis.
  • Knowledge of modern deployment strategies (e.g., blue‑green deployments, canary releases) and resiliency patterns (circuit breakers, retry mechanisms, etc.).
Preferred Qualifications
  • Experience with distributed systems.
  • Experience with statistical analysis applied to metrics.
  • Familiarity with high‑performance, low‑latency systems.
  • Experience as on‑call engineer.
  • Hands‑on experience running Chaos Engineering drills and initiatives.

Zeta considers applicants for employment without regard to, and does not discriminate on the basis of an individual’s sex, race, color, religion, age, disability, status as a veteran, or national or ethnic origin, or any other basis protected by applicable federal, state or local law; nor does Zeta discriminate on the basis of sexual orientation or gender identity or expression.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Lead Site Reliability Engineer
Lead Site Reliability Engineer

Sierra Ventures • Bengaluru

On-site
INR 3,500,000 - 5,500,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Five9 • Bengaluru

On-site
INR 4,000,000 - 7,000,000
Staff Site Reliability Engineer
Staff Site Reliability Engineer

Stryker Group • Gurugram District

On-site
INR 1,200,000 - 1,800,000
Staff Site Reliability Engineer
Staff Site Reliability Engineer

PowerToFly • Gurgaon

On-site
INR 1,800,000 - 3,000,000
Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

Visa Consolidated Support Services India • Bengaluru

On-site
INR 1,500,000 - 2,500,000
Site Reliability Engineer Lead
Site Reliability Engineer Lead

Synechron • Bengaluru, Hyderabad

Hybrid
INR 4,200,000 - 6,300,000
Lead SRE
Lead SRE

Cvent, Inc. • Gurugram District

On-site
INR 4,000,000 - 8,000,000
Senior Cloud Site Reliability Engineer
Senior Cloud Site Reliability Engineer

Augusta Infotech • Bengaluru

Hybrid
INR 1,500,000 - 2,500,000
Site Reliability Engineer
Site Reliability Engineer

Spot Your Leaders & Consulting • Pune District

On-site
INR 2,500,000 - 4,000,000
Lead SRE
Lead SRE

Cvent, Inc. • India

On-site
INR 2,500,000 - 4,500,000