Campus SRE: Reliability Leader for Fleet-Scale Incidents

SpaceXAI

Southaven (MS)

On-site

USD 110,000 - 165,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

SpaceXAI is seeking a Site Reliability Engineer focused on campus reliability to design and govern the systems that the campus watches and trusts. You will coordinate cross-discipline SEVs and build guardrails that shrink incident size.

You will connect compute, network, storage, power, and cooling, guiding calm incident leadership and fleet-scale observability across software and facilities. Ideal candidates bring power plant, data center, or large facilities experience, proven incident-command

Qualifications

  • Bachelor's degree or equivalent experience in a technical discipline.
  • 5+ years of site reliability, systems engineering, plant operations, or large-scale production operations.
  • Proven incident command experience and calm technical leadership on a bridge.
  • Experience designing monitoring and observability at fleet or campus scale.
  • Experience across at least two: compute, network, storage, power, and cooling / facilities telemetry.
  • Experience writing and operating playbooks/runbooks for 24/7 operations.

Responsibilities

  • Own monitoring architecture and signal quality, define what to alert and suppress.
  • Provide SEV command support and bridge coordination with the NOC.
  • Run blameless postmortems and drive corrective actions to closure.
  • Lead cross-functional reliability projects across compute, network, storage, and facilities.
  • Build and maintain playbooks and runbooks; manage dependency maps with NOC.
  • Define error budgets and availability objectives at campus/service boundaries.
  • Participate in on-call rotations and incident response for SEV-class events.

Skills

Python
Bash
C
C++
Java
Go
Rust

Education

Bachelor's degree in Systems Engineering, Computer Science, Electrical Engineering, or a related field

Job description

SpaceXAI is seeking a Site Reliability Engineer focused on campus reliability to design and govern the systems that the campus watches and trusts. You will coordinate cross-discipline SEVs and build guardrails that shrink incident size.

You will connect compute, network, storage, power, and cooling, guiding calm incident leadership and fleet-scale observability across software and facilities. Ideal candidates bring power plant, data center, or large facilities experience, proven incident-command

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Campus Reliability SRE: Incident Command & Observability
Campus Reliability SRE: Incident Command & Observability

Spacex • Memphis (TN), Northern (KY)

Hybrid
USD 140,000 - 190,000
Campus SRE: Incident Leadership & Fleet Reliability
Campus SRE: Incident Leadership & Fleet Reliability

SpaceXAI • Memphis (TN)

On-site
USD 120,000 - 180,000
Site Reliability Engineer
Site Reliability Engineer

Spacex • Memphis (TN), Northern (KY)

On-site
USD 140,000 - 190,000
Site Reliability Engineer - Memphis
Site Reliability Engineer - Memphis

Maven Ventures • Southaven (MS)

On-site
USD 130,000 - 180,000
Site Reliability Engineer, Data Center Infra — Onsite
Site Reliability Engineer, Data Center Infra — Onsite

Spacex • Bastrop (TX)

On-site
USD 140,000 - 200,000
Site Reliability Engineer — Campus Reliability Lead
Site Reliability Engineer — Campus Reliability Lead

Maven Ventures • Southaven (MS)

On-site
USD 130,000 - 180,000
Site Reliability Engineer - Memphis
Site Reliability Engineer - Memphis

SpaceXAI • Memphis (TN)

On-site
USD 120,000 - 180,000
Site Reliability Engineer - Memphis
Site Reliability Engineer - Memphis

SpaceXAI • Southaven (MS)

On-site
USD 110,000 - 165,000
Senior SRE – Platform Infrastructure (Onsite)
Senior SRE – Platform Infrastructure (Onsite)

Spacex • Bastrop (TX), Northern (KY)

Hybrid
USD 170,000 - 230,000
Production Site Reliability Engineer—Mission-Critical Infra
Production Site Reliability Engineer—Mission-Critical Infra

Starlink • Hawthorne (CA)

On-site
USD 125,000 - 195,000
Stock options
401(k) plan
Medical, vision, and dental insurance