Senior SRE: Kubernetes at Scale & Automation

National Geographic

Boston (MA)

On-site

USD 128,000 - 160,000

Full time

7 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

DraftKings is seeking a Senior Site Reliability Engineer in Boston to build and scale our Kubernetes infrastructure across public clouds and on-prem environments. You will design automation-first solutions, implement GitOps delivery with Rancher Fleet, Flux, and Helm, and own scaling and capacity strategies with Karpenter, HPA, and KEDA.

You will develop reliable tooling using Go and Python, monitor reliability with Datadog, and contribute to architectural discussions while participating in

Qualifications

  • Bachelor’s Degree in Computer Science or a related field, or equivalent education, experience, and training.
  • At least 4 years of experience managing distributed cloud and on-premise environments at scale, including strong hands-on experience with Amazon Web Services; experience with Google Cloud Platform, vSphere, or Nutanix is a plus.
  • Deep expertise in Kubernetes and container orchestration, with experience designing, scaling, and troubleshooting complex workloads.
  • Strong software development experience using languages such as Go and Python to build automation and infrastructure tooling.
  • Working knowledge of networking and Linux-based systems, including container runtimes such as Docker and containerd, packet-level debugging, and kernel troubleshooting.
  • Experience with Infrastructure as Code and configuration management tools to build scalable, consistent, and repeatable infrastructure.

Responsibilities

  • Drive stability, performance, and scalability across our global compute platform spanning multiple public clouds and on-premise environments.
  • Build self-healing, fault-tolerant infrastructure and internal tooling that automates repetitive operational work and reduces toil for Platform and Application teams.
  • Operate and evolve our GitOps delivery model, using Rancher Fleet, Flux, and Helm to deploy core Kubernetes services and application workloads consistently and reliably.
  • Own Kubernetes scaling and capacity strategies using technologies including Karpenter, Horizontal Pod Autoscaler (HPA), Kubernetes Event-Driven Autoscaling (KEDA), and predictive scaling based on event and calendar data.
  • Define and monitor service-level objectives and reliability metrics for platform components using Datadog and our logging pipeline.
  • Strengthen our engineering practices by sharing knowledge, contributing to architectural and design discussions, and participating in an on-call rotation.

Skills

Kubernetes
Go
Python
AWS
Linux
Networking
Container runtimes
IaC tools

Education

Bachelor's degree in Computer Science

Tools

Docker
containerd
Terraform

Job description

DraftKings is seeking a Senior Site Reliability Engineer in Boston to build and scale our Kubernetes infrastructure across public clouds and on-prem environments. You will design automation-first solutions, implement GitOps delivery with Rancher Fleet, Flux, and Helm, and own scaling and capacity strategies with Karpenter, HPA, and KEDA.

You will develop reliable tooling using Go and Python, monitor reliability with Datadog, and contribute to architectural discussions while participating in

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior SRE: Kubernetes, Cloud Reliability & Automation
Senior SRE: Kubernetes, Cloud Reliability & Automation

DraftKings • Boston (MA)

On-site
USD 128,000 - 160,000
Senior Site Reliability Engineer: Kubernetes Automation
Senior Site Reliability Engineer: Kubernetes Automation

DraftKings Inc. • Boston (MA)

On-site
USD 128,000 - 160,000
Senior SRE: Remote, Scalable Kubernetes & Automation Lead
Senior SRE: Remote, Scalable Kubernetes & Automation Lead

Camunda • Boston (MA)

Remote
USD 150,000 - 242,000
Remote & Flexible
Annual Kickoff & travel
Health & Wellbeing
+2
Senior SRE - Kubernetes, Observability & Automation (Remote)
Senior SRE - Kubernetes, Observability & Automation (Remote)

Camunda • Atlanta (GA)

Remote
USD 150,000 - 242,000
Remote work
Annual company events
Health & wellbeing
+2
Senior Cloud SRE: Kubernetes, GitOps & Automation
Senior Cloud SRE: Kubernetes, GitOps & Automation

Pinterest • San Francisco (CA)

On-site
USD 140,000 - 288,000
Senior SRE – Cloud-native, Kubernetes & CI/CD
Senior SRE – Cloud-native, Kubernetes & CI/CD

Pinterest • San Francisco (CA)

On-site
USD 139,764 - 287,749
Equity
Competitive salary
Senior SRE, DevEx Platform & Kubernetes Automation
Senior SRE, DevEx Platform & Kubernetes Automation

Chainlink Labs • Las Vegas (NV)

On-site
USD 150,000 - 180,000
Lead Site Reliability Engineer: Drive SLOs & Equity
Lead Site Reliability Engineer: Drive SLOs & Equity

DraftKings Inc. • Boston (MA)

On-site
USD 148,000 - 185,000
Senior SRE: Cloud Reliability, Terraform & Kubernetes (Remote)
Senior SRE: Cloud Reliability, Terraform & Kubernetes (Remote)

Motion Recruitment • Chicago (IL)

On-site
USD 140,000 - 190,000
Staff SRE: Scale, Observability & Kubernetes
Staff SRE: Scale, Observability & Kubernetes

Replit • Foster City (CA)

On-site
USD 180,000 - 260,000
Competitive Salary & Equity
401(k) with 4% match
Health, Dental, Vision and Life Ins.
+2