Site Reliability Engineer — Campus Reliability Lead

Pantera Capital

Southaven (MS)

On-site

USD 130,000 - 180,000

Full time

7 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

SpaceXAI is seeking a Site Reliability Engineer focused on campus reliability to design monitoring, lead incident command, and align compute, network, storage, power, and cooling. You will own playbooks, drive blameless postmortems, define error budgets, and collaborate with NOC and facilities teams.

The role requires 5+ years in SRE or related fields, hands-on leadership, and strong scripting in Python/Bash with experience in at least one systems language.

Qualifications

  • Bachelor's degree in a related field is required.
  • 5+ years of site reliability, systems engineering, or large-scale production operations experience.
  • Proven large-scale incident command experience and calm technical leadership.
  • Demonstrated monitoring and observability design at fleet or campus scale.
  • Experience across at least two of: compute, network, storage, power, and cooling.
  • Experience writing and operating playbooks with 24/7 operations or NOC partners.
  • Proficiency in scripting (Python, Bash) and at least one systems language.
  • Excellent problem-solving with data-driven reliability approach.
  • Ability to collaborate with cross-functional teams including NOC and data center operations.

Responsibilities

  • Own monitoring architecture and signal quality; manage alerting strategy.
  • Provide SEV command support and incident leadership.
  • Run blameless postmortems and drive corrective actions to closure.
  • Lead cross-functional reliability projects across compute, network, storage, and facilities.
  • Develop and maintain runbooks; coordinate with NOC for operations.
  • Define error budgets and availability objectives across campus/service boundaries.
  • Participate in on-call rotations and incident response for campus data center events in Memphis/Southaven.

Skills

Python
Bash
C/C++
Java
Go
Rust

Education

Bachelor's degree in Systems Engineering / CS / EE

Job description

SpaceXAI is seeking a Site Reliability Engineer focused on campus reliability to design monitoring, lead incident command, and align compute, network, storage, power, and cooling. You will own playbooks, drive blameless postmortems, define error budgets, and collaborate with NOC and facilities teams.

The role requires 5+ years in SRE or related fields, hands-on leadership, and strong scripting in Python/Bash with experience in at least one systems language.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Campus Reliability SRE & Incident Commander
Campus Reliability SRE & Incident Commander

Socket.dev • Southaven (MS)

On-site
USD 120,000 - 190,000
Campus SRE: Reliability Lead for Fleet & Incidents
Campus SRE: Reliability Lead for Fleet & Incidents

Pantera Capital • Southaven (MS)

On-site
USD 120,000 - 180,000
Campus Reliability SRE Lead: Incident Command & Playbooks
Campus Reliability SRE Lead: Incident Command & Playbooks

SpaceXAI • Southaven (MS)

On-site
USD 120,000 - 170,000
Site Reliability Engineer
Site Reliability Engineer

SpaceXAI • Southaven (MS)

On-site
USD 120,000 - 170,000
Site Reliability Engineer
Site Reliability Engineer

Socket.dev • Southaven (MS)

On-site
USD 120,000 - 190,000
Site Reliability Engineer
Site Reliability Engineer

Pantera Capital • Southaven (MS)

On-site
USD 120,000 - 180,000
Site Reliability Engineer
Site Reliability Engineer

Spacex • Memphis (TN), Northern (KY)

Hybrid
USD 140,000 - 190,000
Campus Reliability SRE: Incident Command & Observability
Campus Reliability SRE: Incident Command & Observability

Spacex • Memphis (TN), Northern (KY)

Hybrid
USD 140,000 - 190,000
Site Reliability Engineer - Memphis
Site Reliability Engineer - Memphis

Pantera Capital • Southaven (MS)

On-site
USD 130,000 - 180,000
Site Reliability Engineer — Mission-Critical Systems
Site Reliability Engineer — Mission-Critical Systems

Iac/interactivecorp • Sacramento (CA)

On-site
USD 120,000 - 170,000
Stock options
Health insurance
Paid vacation