Campus Reliability SRE Lead: Incident Command & Playbooks

SpaceXAI

Southaven (MS)

On-site

USD 120,000 - 170,000

Full time

9 days ago
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

SpaceXAI is seeking a Site Reliability Engineer focused on campus reliability in the Memphis/Southaven data center campus. You will design monitoring, lead incident command, and drive cross-discipline reliability work across compute, network, storage, power, and cooling.

You will own runbooks, run game days, define error budgets, and guide blameless postmortems with a strong emphasis on data-driven decisions and cross-functional collaboration with NOC, data center operations, and infrastructure

Qualifications

  • Bachelor’s degree in Systems Engineering, CS, Electrical Engineering, or related field (or equivalent experience).
  • 5+ years in site reliability, systems engineering, or large-scale production ops, HPC or data centers.
  • Proven large-scale incident command experience and calm technical leadership on a bridge.
  • Demonstrated monitoring and observability design at fleet or campus scale, including alert hygiene.
  • Experience across at least two: compute, network, storage, power, and cooling / facilities telemetry.
  • Experience writing and operating playbooks or runbooks with 24/7 operations or NOC.
  • Proficiency in scripting (Python, Bash) for automation and analysis, plus at least one systems language (C, C++, Java, Go, Rust).
  • Excellent problem-solving with a data-driven reliability mindset.
  • Ability to work with cross-functional teams, including NOC, data center ops, and infra engineering.

Responsibilities

  • Own monitoring architecture and signal quality; manage alerts and noise.
  • Provide SEV command support: incident leadership, bridge with NOC, and timeline hygiene.
  • Run blameless postmortems and drive corrective actions to closure.
  • Lead cross-functional reliability projects spanning compute, network, storage, and facility signal boundaries.
  • Build and maintain runbooks, run game days, and update dependency maps with NOC.
  • Define error budgets and availability objectives at campus and service boundaries.
  • Participate in on-call rotations and incident response for SEV-class events in the Memphis / Southaven campus.

Skills

Python
Bash
Go
Rust
C
C++
Java

Education

Bachelor's degree in Systems Engineering / CS / EE or related field

Job description

SpaceXAI is seeking a Site Reliability Engineer focused on campus reliability in the Memphis/Southaven data center campus. You will design monitoring, lead incident command, and drive cross-discipline reliability work across compute, network, storage, power, and cooling.

You will own runbooks, run game days, define error budgets, and guide blameless postmortems with a strong emphasis on data-driven decisions and cross-functional collaboration with NOC, data center operations, and infrastructure

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Campus Reliability SRE & Incident Commander
Campus Reliability SRE & Incident Commander

Socket.dev • Southaven (MS)

On-site
USD 120,000 - 190,000
Campus SRE & Data-Center Reliability Lead
Campus SRE & Data-Center Reliability Lead

SpaceXAI • Southaven (MS)

On-site
USD 120,000 - 180,000
Campus Reliability SRE: Incident Command & Observability
Campus Reliability SRE: Incident Command & Observability

Spacex • Memphis (TN), Northern (KY)

Hybrid
USD 140,000 - 190,000
Site Reliability Engineer — Campus Reliability Lead
Site Reliability Engineer — Campus Reliability Lead

Pantera Capital • Southaven (MS)

On-site
USD 130,000 - 180,000
Site Reliability Engineer - Memphis
Site Reliability Engineer - Memphis

Pantera Capital • Southaven (MS)

On-site
USD 130,000 - 180,000
Site Reliability Engineer - Memphis
Site Reliability Engineer - Memphis

SpaceXAI • Southaven (MS)

On-site
USD 120,000 - 180,000
Site Reliability Engineer
Site Reliability Engineer

SpaceXAI • Southaven (MS)

On-site
USD 120,000 - 170,000
Site Reliability Engineer
Site Reliability Engineer

Spacex • Memphis (TN), Northern (KY)

Hybrid
USD 140,000 - 190,000
Site Reliability Engineer
Site Reliability Engineer

Socket.dev • Southaven (MS)

On-site
USD 120,000 - 190,000
Hardware SRE — Firmware & Datacenter Reliability
Hardware SRE — Firmware & Datacenter Reliability

SpaceXAI • Memphis (TN)

On-site
USD 110,000 - 160,000