Lead Site Reliability Engineer

National Geographic

Boston (MA)

On-site

USD 148,000 - 185,000

Full time

12 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Bonus
Equity
Benefits

Job summary

DraftKings seeks a Lead Site Reliability Engineer in Boston to set the reliability standard across our Infrastructure Engineering organization. You will define how we measure reliability, establish SLOs and error budgets, and translate telemetry into actionable insights for leadership and product teams.

As an individual contributor, you will lead technical guidance, mentor peers, and align infrastructure performance with user experiences.

Qualifications

  • Bachelor's degree in Computer Science or related field, or equivalent education and training.
  • 7+ years in Site Reliability Engineering, with hands-on experience defining and operating SLIs/SLOs and error budgets at scale.
  • Deep experience with observability platforms like Datadog, building dashboards and alerts from metrics/logs.
  • Ability to map reliability objectives to application/platform outcomes and end-user impact.
  • Strong knowledge of distributed systems and failure modes affecting reliability measurement.

Responsibilities

  • Lead and mature the SLO development process across Infrastructure Engineering with clear frameworks and targets.
  • Collaborate with engineering teams to design and implement SLOs, correlating user journeys to metrics, targets, and error budgets.
  • Review and refine reliability objectives to stay aligned with customer impact and business priorities.
  • Tie infrastructure reliability targets to end-user experiences and dependencies.
  • Build reporting processes and tools to provide visibility of reliability across critical components.
  • Influence reliability practices through mentoring, design reviews, and cross-team collaboration.
  • Differentiate meaningful degradation from downtime using thoughtful measurement in distributed systems.

Skills

Service Level Objectives
Datadog
Observability
Distributed systems
AWS/Kubernetes
Cross-functional influence

Education

Bachelor's Degree in CS or related field

Tools

Datadog
Kubernetes
AWS

Job description

At DraftKings, AI is becoming an integral part of both our present and future, powering how work gets done today, guiding smarter decisions, and sparking bold ideas. It’s transforming how we enhance customer experiences, streamline operations, and unlock new possibilities. Our teams are energized by innovation and readily embrace emerging technology. We’re not waiting for the future to arrive. We’re shaping it, one bold step at a time. To those who see AI as a driver of progress, come build the future together.

The Crown Is Yours As a Lead Site Reliability Engineer, you’ll set the reliability standard across our Infrastructure Engineering organization. You’ll define how we measure reliability for critical services, partnering with engineering teams to build and refine Service Level Objectives that connect infrastructure performance to the experiences our platforms support. You’ll turn complex telemetry into clear, actionable insights that help teams and senior leaders make better decisions about reliability, risk, and priorities. As an individual contributor, you’ll lead through technical expertise and influence, shaping a consistent reliability practice across the organization.

What you’ll do as a Lead Site Reliability Engineer
  • Lead and mature the Service Level Objective development process across Infrastructure Engineering, establishing clear frameworks and standards for setting meaningful reliability targets.
  • Partner with engineering teams to design and implement Service Level Objectives, beginning with the critical user journeys each service supports and translating them into measurable indicators, targets, and error budget policies.
  • Review and refine existing reliability objectives to keep them aligned with changing customer impact, technical dependencies, and business priorities.
  • Connect infrastructure reliability targets to the application and platform experiences they support, making dependencies and their impact on end-user experience clear and measurable.
  • Build reporting processes and tooling that provide a clear view of reliability across critical components, translating technical signals into actionable insights for Senior Managers and Directors.
  • Influence reliability practices across teams through technical guidance, design reviews, mentoring, and collaboration with engineering partners.
  • Help teams distinguish meaningful service degradation from true downtime by applying thoughtful measurement strategies to complex distributed systems.
What you’ll bring
  • A Bachelor’s Degree in Computer Science or a related field, or equivalent relevant education, experience, and training.
  • At least 7 years of experience in Site Reliability Engineering, including hands‑on experience defining and operationalizing Service Level Objectives, Service Level Indicators, and error budgets at scale.
  • Deep experience with observability platforms such as Datadog, including building dashboards, monitors, and reporting from metrics and logging pipelines.
  • Experience connecting infrastructure‑level reliability objectives to application or platform‑level outcomes and evaluating how technical dependencies affect end‑user experience.
  • Strong knowledge of distributed systems and the failure modes that can make reliability measurement complex, with the ability to assess what technical signals truly represent.
  • Proven ability to influence across engineering teams, translate reliability concepts for technical and non‑technical audiences, and drive alignment without direct authority.
  • Excellent written and verbal communication skills, including experience developing and presenting reliability reporting to Senior Managers, Directors, and cross‑functional stakeholders.
  • Working knowledge of cloud and infrastructure environments such as Amazon Web Services, Kubernetes, and on‑premise systems, with the technical depth to partner effectively with the teams operating them.
Join Our Team

We’re a publicly traded (NASDAQ: DKNG) technology company headquartered in Boston. As a regulated gaming company, you may be required to obtain a gaming license issued by the appropriate state agency as a condition of employment. Don’t worry, we’ll guide you through the process if this is relevant to your role.

The US base salary range for this full-time position is 148,000.00 USD - 185,000.00 USD, plus bonus, equity, and benefits as applicable. Our ranges are determined by role, level, and location. The compensation information displayed on each job posting reflects the range for new hire pay rates for the position across all US locations. Within the range, individual pay is determined by work location and additional factors, including job-related skills, experience, and relevant education or training. Your recruiter can share more about the specific pay range and how that was determined during the hiring process.

It is unlawful in Massachusetts to require or administer a lie detector test as a condition of employment or continued employment. An employer who violates this law shall be subject to criminal penalties and civil liability.

DraftKings Inc. (Nasdaq: DKNG) is a digital sports entertainment and gaming company. It’s simple, at DraftKings, we believe life’s more fun with skin in the game. For that reason, we’re committed to responsibly creating the world’s favorite games and betting experiences. Headquartered in Boston, with offices around the globe, we believe we can continue to define what it means to be a technology company in sports entertainment. We love what we do, and think you will too.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Lead Site Reliability Engineer
Lead Site Reliability Engineer

DraftKings • Boston (MA), Northern (KY)

Hybrid
USD 148,000 - 185,000
Lead Site Reliability Engineer
Lead Site Reliability Engineer

DraftKings Inc. • Boston (MA)

On-site
USD 148,000 - 185,000
Senior Site Reliability Engineer, Infrastructure
Senior Site Reliability Engineer, Infrastructure

DraftKings • Boston (MA)

On-site
USD 128,000 - 160,000
Senior Lead Database Reliability Engineer
Senior Lead Database Reliability Engineer

DraftKings Inc. • Massachusetts

On-site
USD 168,000 - 210,000
AI-Driven Database Reliability Engineer
AI-Driven Database Reliability Engineer

DraftKings Inc. • Boston (MA)

On-site
USD 112,000 - 140,000
Database Reliability Engineer
Database Reliability Engineer

DraftKings Inc. • Boston (MA)

On-site
USD 112,000 - 140,000
Director, Leadership Effectiveness
Director, Leadership Effectiveness

National Geographic • Boston (MA)

On-site
USD 172,000 - 215,000
Data Engineering Manager, Customer
Data Engineering Manager, Customer

National Geographic • Boston (MA)

On-site
USD 153,000 - 192,000
Bonus
Equity
Benefits
Senior Technical Compliance Analyst
Senior Technical Compliance Analyst

National Geographic • Boston (MA)

On-site
USD 111,000 - 139,000
VIP Lifecycle Associate, Predictions
VIP Lifecycle Associate, Predictions

National Geographic • United States

Remote
USD 54,000 - 68,000
Bonus
Equity
Benefits