SRE Lead for 24x7 AI Platform Reliability

Graphcore

Austin (TX)

On-site

USD 170,000 - 230,000

Full time

2 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Flexible working
Medical/Dental/Vision coverage
401(k) retirement plan
FSAs/HSAs
Disability and life insurance
Commuter benefits

Job summary

Graphcore seeks an experienced Site Reliability Engineering leader to build and lead a new SRE organization responsible for the production operation of a rapidly scaling AI supercomputing platform. This role blends hands-on engineering with organizational leadership to ensure 24x7x365 availability and reliability across compute, networking, storage, and orchestration layers.

The SRE Manager will hire, mentor, and develop the team, define operating models, and drive automation and reliability

Qualifications

  • Significant experience leading or building SRE/Production engineering teams for large-scale, highly available production infrastructure.
  • Experience taking a new or evolving platform through production readiness, launch and stabilization.
  • Strong incident leadership, including managing high-severity, multi-team incidents.

Responsibilities

  • Build the SRE organization from formation to 24x7x365 production operations.
  • Define the operating model, staffing, escalation, on-call, change management and readiness.
  • Embed with platform engineering to ensure reliability is designed in from the start.
  • Lead post-incident reviews and drive systemic improvements and automation.

Skills

SRE leadership
Production engineering
Incident management
Observability
Automation
Linux basics

Tools

Kubernetes
Slurm
Datacenter operations

Job description

Graphcore seeks an experienced Site Reliability Engineering leader to build and lead a new SRE organization responsible for the production operation of a rapidly scaling AI supercomputing platform. This role blends hands-on engineering with organizational leadership to ensure 24x7x365 availability and reliability across compute, networking, storage, and orchestration layers.

The SRE Manager will hire, mentor, and develop the team, define operating models, and drive automation and reliability

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

SRE Lead for 24x7 AI Platform Reliability
SRE Lead for 24x7 AI Platform Reliability

Graphcore • Austin (TX)

Hybrid
USD 170,000 - 260,000
Medical insurance
Dental insurance
Vision insurance
+2
Remote SRE Manager: Lead AI-Driven Reliability & Cloud Ops
Remote SRE Manager: Lead AI-Driven Reliability & Cloud Ops

Arcoro Holdings Corp • Phoenix (AZ), Northern (KY)

Hybrid
USD 200,000 - 220,000
Remote Work
401(k) with Company match
Flexible PTO and Company-paid holidays
Senior SRE Lead: Scale Reliability & AI Ops
Senior SRE Lead: Scale Reliability & AI Ops

NVIDIA Gruppe • Santa Clara (CA)

Hybrid
USD 168,000 - 334,000
Equity
Benefits
SRE Director: Build Global Reliability & Infra
SRE Director: Build Global Reliability & Infra

d-Matrix • Santa Clara (CA)

On-site
USD 180,000 - 240,000
SRE Operations Lead — AI-Driven Cloud & CI/CD
SRE Operations Lead — AI-Driven Cloud & CI/CD

SFE • Charlotte (NC)

On-site
USD 140,000 - 180,000
AI Platform & SRE Engineering Lead
AI Platform & SRE Engineering Lead

Jobs in JS • Santa Clara (CA)

On-site
USD 208,000 - 334,000
Remote SRE Leader: Scale Infra & Reliability
Remote SRE Leader: Scale Infra & Reliability

Invoca • San Francisco (CA)

Remote
USD 190,000 - 250,000
Health Benefits
Mental Wellbeing
Wellness Subsidy
+5
Global SRE Leader: Platform Reliability & Automation
Global SRE Leader: Platform Reliability & Automation

Broadridge Financial Solutions • New York (NY)

Hybrid
USD 235,000 - 250,000
Remote SRE Lead: Drive AI-Driven Reliability
Remote SRE Lead: Drive AI-Driven Reliability

Fingerprint • Chicago (IL)

Remote
USD 177,000 - 240,000
AI Infrastructure SRE Manager – Scale & Reliability
AI Infrastructure SRE Manager – Scale & Reliability

Google • San Jose (CA)

On-site
USD 207,000 - 300,000