Senior Site Reliability Engineer, Data & Analytics

Blizzard Entertainment

Irvine (CA)

Hybrid

USD 101,000 - 186,754

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Medical, dental, and vision insurance
401(k) with company match
Paid holidays and vacation
Mental health and wellbeing programs

Job summary

Blizzard Entertainment is seeking a Senior Site Reliability Engineer for its Data & Analytics team. You will enhance the reliability and performance of large-scale data systems while building automation tools.

The ideal candidate should have expertise in distributed systems, Kubernetes, and must be proficient in Linux and automation tools. The position offers competitive benefits and a flexible working environment, allowing for hybrid working options.

Qualifications

  • Experience operating reliable, distributed systems in SRE or similar roles.
  • Strong knowledge of Linux, containers, and cloud infrastructure.
  • Experience with CI/CD or GitOps systems.

Responsibilities

  • Participate in on-call rotation and drive incidents to resolution.
  • Design and build automation and operational tooling.
  • Build and evolve centralized platform services.

Skills

Distributed systems operation
Data analytics experience
Kubernetes knowledge
Infrastructure-as-code
Automation tools
Observability practices
Strong communication skills

Tools

Terraform
CI/CD tools
GitHub Actions
Prometheus

Job description

Senior Site Reliability Engineer, Data & Analytics

The Senior Site Reliability Engineer role is on our Data & Analytics team, partnering with data, analytics, ML, and platform engineering to improve the reliability, scalability, and performance of large‑scale data platforms, analytics pipelines, ML training pipelines, and inference services.

In addition to core SRE responsibilities, this role will build operational and automation tooling that reduces toil, speeds up issue resolution, and improves engineering velocity. This includes contributing to internal platform services such as shared tooling, data integrations, and access‑control patterns used across Blizzard.

The ideal candidate is a production‑minded SRE or platform engineer who is comfortable operating critical systems, writing software, and building tools that improve engineering efficiency without compromising reliability.

This role is open to candidates based in Irvine, CA or Albany, NY (hybrid or on‑site), as well as fully remote candidates.

Responsibilities
  • Participate in an on‑call rotation and drive incidents to resolution
  • Lead blameless postmortems and identify systemic reliability improvements
  • Partner with data, ML, and platform teams to improve batch, streaming, training, and inference workloads
  • Support ML training pipelines and inference services, including GPU workloads
  • Help define how data and ML services run on Kubernetes
  • Design and build automation and operational tooling (e.g., workflows, diagnostic tooling, runbooks) to reduce on‑call burden
  • Build and evolve centralized platform services, including shared tooling, data integrations, and access controls
  • Diagnose and resolve reliability, performance, and cost issues across distributed systems
  • Champion automation, documentation, and practices that reduce toil
  • Maintain infrastructure using Terraform and infrastructure‑as‑code principles
  • Improve CI/CD and GitOps workflows (Jenkins, GitHub Actions, ArgoCD)
  • Operate and improve containerized services on Kubernetes
  • Define and measure reliability using SLIs, SLOs, and error budgets
  • Run load tests, capacity modeling, and production validation
  • Build internal tools and paved paths that help teams operate safely and efficiently
Minimum Requirements
  • Experience operating reliable, distributed systems in SRE, platform, or similar roles
  • Experience with data, analytics, ML, or large‑scale distributed workloads
  • Strong knowledge of Linux, containers, Kubernetes, and cloud infrastructure
  • Experience building automation or internal tools (Python, Go, shell, etc.)
  • Experience with infrastructure‑as‑code (e.g., Terraform)
  • Experience with CI/CD or GitOps systems (e.g., Jenkins, GitHub Actions, ArgoCD)
  • Familiarity with observability (metrics, logs, traces, alerting, incident response)
  • Solid understanding of SRE concepts (SLIs, SLOs, error budgets, postmortems)
  • Experience using modern development and automation practices to improve reliability and efficiency
  • Experience building internal tooling, automation, or developer productivity systems
  • Strong communication skills with technical and cross‑functional partners
Bonus Points
  • Experience with data and ML systems (training pipelines, model serving, GPU workloads)
  • Experience with distributed systems and messaging (Kafka, Pub/Sub)
  • Experience working in Kubernetes‑based environments
  • Familiarity with observability tools (Prometheus, Grafana)
  • Experience operating systems in cloud environments (GCP, AWS)
Benefits
  • Medical, dental, vision, health savings account or health reimbursement account, healthcare spending accounts, dependent care spending accounts, life and AD&D insurance, disability insurance
  • 401(k) with company match, tuition reimbursement, charitable donation matching
  • Paid holidays and vacation, paid sick time, floating holidays, compassion and bereavement leaves, parental leave
  • Mental health & wellbeing programs, fitness programs, free and discounted games, and a variety of other voluntary benefit programs like supplemental life & disability, legal service, ID protection, rental insurance, and others
  • Relocation assistance if the company requires geographic mobility

In the U.S., the standard base pay range for this role is $101,000.00 – $186,754.00 annually. Compensation is based on experience, performance, and location.

We are an equal opportunity employer and value diversity at our company. We do not discriminate on the basis of race, religion, color, national origin, gender, sexual orientation, gender identity, age, marital status, veteran status, or disability status, among other characteristics.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Associate Site Reliability Engineer
Associate Site Reliability Engineer

Activision Blizzard • Santa Monica (CA)

On-site
Medical, dental, vision insurance
401(k) with company match
Paid holidays and vacation
+3
Associate Site Reliability Engineer
Associate Site Reliability Engineer

Activision • Santa Monica (CA)

On-site
USD 75,000 - 95,000
Medical, dental, and vision insurance
401(k) with company match
Paid holidays and vacation
+2
Senior SRE - Data & Analytics (Remote-friendly)
Senior SRE - Data & Analytics (Remote-friendly)

BLIZZARD ENTERTAINMENT, INC. • Irvine (CA)

Hybrid
USD 103,000 - 190,000
Medical, dental, vision benefits
401(k) with Company match
Paid holidays and vacation
Senior Data & Analytics SRE — Remote
Senior Data & Analytics SRE — Remote

Blizzard Entertainment • Irvine (CA)

Hybrid
USD 101,000 - 187,000
Medical, dental, and vision insurance
401(k) with company match
Paid holidays and vacation
+1
Senior Site Reliability Engineer, Data & Analytics
Senior Site Reliability Engineer, Data & Analytics

BLIZZARD ENTERTAINMENT, INC. • Irvine (CA)

Hybrid
USD 103,000 - 190,000
Medical, dental, vision benefits
401(k) with Company match
Paid holidays and vacation
Lead, Site Reliability Engineer
Lead, Site Reliability Engineer

CardWorks • Pittsburgh

Hybrid
USD 146,000 - 163,000
Competitive Pay
Medical, Dental, and Vision Benefits
401(k) Plan with Company Match
+1
Senior Site Reliability Engineer - Data Infrastructure (San Jose)
Senior Site Reliability Engineer - Data Infrastructure (San Jose)

ByteDance • San Jose (CA)

On-site
USD 212,800 - 387,600
Senior Site Reliability Engineer -AI Infrastructure Operations
Senior Site Reliability Engineer -AI Infrastructure Operations

Nscale • San Francisco (CA), Seattle (WA), Houston (TX)

On-site
USD 170,000 - 265,000
Equity
Ownership from start
Flexible schedule
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Storm2 • Scottsdale (AZ)

Hybrid
USD 140,000 - 150,000
Competitive healthcare, dental, and vision coverage
401(k) with company match
Generous PTO and paid holidays
+1
Senior Site Reliability Engineer II
Senior Site Reliability Engineer II

RELX • New York (NY)

Hybrid
USD 104,000 - 175,000
Competitive salary
Hybrid or remote options
Career growth in SRE and DevOps
+1