Campus Reliability SRE: Incident Command & Observability

Spacex

Memphis, Northern (TN, KY)

Hybrid

USD 140,000 - 190,000

Full time

6 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

SpaceXAI is seeking a Site Reliability Engineer focused on campus reliability in Memphis. You will design what the campus watches and trusts, lead cross-disciplinary SEVs, and build guardrails to reduce incidents.

You are the bridge across compute, network, storage, power, and cooling while guiding reliability efforts. Ideal candidates have strong incident leadership and fleet-scale observability skills, with a track record in observability design and on-call operations for a data center

Qualifications

  • Bachelor's degree in Systems Engineering, Computer Science, Electrical Engineering, or related field (or equivalent).
  • 5+ years in site reliability, systems engineering, or large-scale production operations, preferably HPC/data centers.
  • Proven large-scale incident command experience and calm technical leadership on a bridge.
  • Demonstrated monitoring and observability design at fleet or campus scale, including alert hygiene, suppression, and signal quality.
  • Experience across at least two of: compute, network, storage, power, and cooling / facilities telemetry.
  • Experience writing and operating playbooks or runbooks with a 24/7 operations or NOC partner.
  • Proficiency in scripting (Python, Bash) for automation and analysis, plus general experience in at least one systems language (C, C++, Java, Go, Rust, or similar). Not required to be expert in all.

Responsibilities

  • Own monitoring architecture and signal quality: what we alert on, suppress, and trust; drive suppression and redesign.
  • Provide SEV command support: technical incident leadership, bridge coordination with the NOC, and timeline hygiene.
  • Run blameless postmortems and drive corrective actions to closed.
  • Lead cross-functional reliability projects spanning compute, network, storage, and facility signal boundaries.
  • Build and maintain runbooks, run game days, and keep dependency maps current; own runbook quality with NOC.
  • Define error budgets and availability objectives at campus and service boundaries as adopted by the business.
  • Participate in on-call rotations and incident response for SEV-class events in the Memphis / Southaven data center campus.

Skills

Python
Bash
C/C++/Java/Go/Rust
Problem solving
Cross-functional collaboration

Education

Bachelor's degree in Systems Engineering, Computer Science, Electrical Engineering, or a related field

Tools

Python
Bash

Job description

SpaceXAI is seeking a Site Reliability Engineer focused on campus reliability in Memphis. You will design what the campus watches and trusts, lead cross-disciplinary SEVs, and build guardrails to reduce incidents.

You are the bridge across compute, network, storage, power, and cooling while guiding reliability efforts. Ideal candidates have strong incident leadership and fleet-scale observability skills, with a track record in observability design and on-call operations for a data center

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Campus Reliability SRE & Incident Commander
Campus Reliability SRE & Incident Commander

Socket.dev • Southaven (MS)

On-site
USD 120,000 - 190,000
Campus SRE & Data-Center Reliability Lead
Campus SRE & Data-Center Reliability Lead

SpaceXAI • Southaven (MS)

On-site
USD 120,000 - 180,000
Campus SRE: Reliability Lead for Fleet & Incidents
Campus SRE: Reliability Lead for Fleet & Incidents

Pantera Capital • Southaven (MS)

On-site
USD 120,000 - 180,000
Campus Reliability SRE Lead: Incident Command & Playbooks
Campus Reliability SRE Lead: Incident Command & Playbooks

SpaceXAI • Southaven (MS)

On-site
USD 120,000 - 170,000
Site Reliability Engineer — Campus Reliability Lead
Site Reliability Engineer — Campus Reliability Lead

Pantera Capital • Southaven (MS)

On-site
USD 130,000 - 180,000
Site Reliability Engineer - Memphis
Site Reliability Engineer - Memphis

Pantera Capital • Southaven (MS)

On-site
USD 130,000 - 180,000
Site Reliability Engineer
Site Reliability Engineer

SpaceXAI • Southaven (MS)

On-site
USD 120,000 - 170,000
Site Reliability Engineer - Memphis
Site Reliability Engineer - Memphis

SpaceXAI • Southaven (MS)

On-site
USD 120,000 - 180,000
Site Reliability Engineer
Site Reliability Engineer

Pantera Capital • Southaven (MS)

On-site
USD 120,000 - 180,000
Site Reliability Engineer
Site Reliability Engineer

Socket.dev • Southaven (MS)

On-site
USD 120,000 - 190,000