Site Reliability Engineer - Memphis

SpaceXAI

Memphis (TN)

On-site

USD 120,000 - 180,000

Full time

40 hours ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

SpaceXAI is seeking a Site Reliability Engineer focused on campus reliability to design what the campus watches and trusts, and to drive reliability across compute, network, storage, power, and cooling. You will lead incident response with calm authority and partner across software and facilities teams in the Memphis campus.

You will own observability, runbooks, and cross-functional initiatives to reduce incident size, while shaping error budgets and service-level objectives for complex

Qualifications

  • Bachelor's degree in Systems Engineering, Computer Science, Electrical Engineering, or related field, or equivalent experience.
  • 5+ years of site reliability, systems engineering, or large-scale production ops, preferably HPC or data centers.
  • Proven large-scale incident command experience and calm technical leadership on a bridge.
  • Experience with fleet- or campus-scale monitoring and observability design, alert hygiene and signal quality.
  • Experience across at least two areas: compute, network, storage, power, or cooling.
  • Experience writing and operating playbooks or runbooks in 24/7 environments.
  • Scripting in Python or Bash for automation; knowledge of a system language (C/C++, Java, Go, Rust).

Responsibilities

  • Own monitoring architecture and signal quality: what we alert on, suppress, and trust.
  • Provide SEV command support: incident leadership and NOC coordination.
  • Run blameless postmortems and drive corrective actions to closure.
  • Lead cross-functional reliability projects spanning compute, network, storage, and facility signals.
  • Build and maintain playbooks, run game days, and map dependencies with the NOC.
  • Define error budgets and availability objectives at campus and service boundaries.
  • Participate in on-call rotations and incident response for SEV-class events in the Memphis / Southaven data center campus.

Skills

SRE leadership
Python
Bash
C/C++/Java/Go/Rust
Data-driven reliability
Incident management

Education

Bachelor's degree in Systems Engineering/CS/EE or related field

Tools

Python
Bash

Job description

SpaceXAI's mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who appreciate challenging themselves and thrive on curiosity. We operate with a flat organizational structure. All employees are expected to be hands-on and to contribute directly to the company’s mission. Leadership is given to those who show initiative and consistently deliver excellence. Work ethic and strong prioritization skills are important. All employees are expected to have strong communication skills. They should be able to concisely and accurately share knowledge with their teammates.

SpaceXAI's mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who appreciate challenging themselves and thrive on curiosity. We operate with a flat organizational structure. All employees are expected to be hands-on and to contribute directly to the company’s mission. Leadership is given to those who show initiative and consistently deliver excellence. Work ethic and strong prioritization skills are important. All employees are expected to have strong communication skills. They should be able to concisely and accurately share knowledge with their teammates.

ABOUT THE ROLE:

As a Site Reliability Engineer focused on campus reliability, you will design what the campus watches and trusts, technically command cross-discipline SEVs, and build the guardrails that make the next incident smaller. You are the connective tissue across compute, network, storage, power, and cooling. This role demands calm incident leadership, fleet-scale observability judgment, and the ability to drive reliability work across software and facility boundaries.

RESPONSIBILITIES:
  • Own monitoring architecture and signal quality: what we alert on, suppress, and trust. Consume NOC noise-disposition feedback to drive suppression and redesign. Treat alert noise as a design failure, not an operator failure.
  • Provide SEV command support: technical incident leadership, bridge coordination with the NOC, and timeline and severity hygiene.
  • Run blameless postmortems and drive corrective actions to closed, not filed.
  • Lead cross-functional reliability projects spanning compute, network, storage, and facility signal boundaries.
  • Build and maintain playbooks, run game days, and keep cross-discipline dependency maps current. Own runbook quality jointly with the NOC (SRE designs; NOC operates and corrects).
  • Define error budgets and availability objectives at campus and service boundaries as adopted by the business.
  • Participate in on-call rotations and incident response for SEV-class events in the Memphis / Southaven data center campus.
BASIC QUALIFICATIONS:
  • Bachelor's degree in Systems Engineering, Computer Science, Electrical Engineering, or a related field (or equivalent experience).
  • 5+ years of experience in site reliability, systems engineering, or large-scale production operations, preferably in high-performance computing or data center environments.
  • Proven large-scale incident command experience and calm technical leadership on a bridge.
  • Demonstrated monitoring and observability design at fleet or campus scale, including alert hygiene, suppression, and signal quality.
  • Experience working across at least two of: compute, network, storage, power, and cooling / facilities telemetry.
  • Experience writing and operating playbooks or runbooks with a 24/7 operations or NOC partner.
  • Proficiency in scripting (Python, Bash) for automation and analysis, plus general experience in at least one systems language (C, C++, Java, Go, Rust, or similar). Not required to be expert in all of them.
  • Excellent problem-solving skills with a data-driven approach to reliability engineering.
  • Ability to work collaboratively with cross-functional teams, including NOC, data center operations, and infrastructure engineering.
PREFERRED SKILLS AND EXPERIENCE:
  • Experience in AI/ML infrastructure or supercomputing environments.
  • Hands-on definition and use of SLOs, SLIs, and error budgets at service or campus boundaries.
  • Experience running game days, dependency mapping, and closed-loop corrective action programs.
  • Familiarity with data center hardware and plant signals (servers, GPUs, networking, power, cooling) in addition to software telemetry.
  • Prior work in a fast-paced startup or tech company like SpaceXAI.

SpaceXAI is an equal opportunity employer. For details on data processing, view our Recruitment Privacy Notice.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer - Memphis
Site Reliability Engineer - Memphis

Pantera Capital • Southaven (MS)

On-site
USD 130,000 - 180,000
Site Reliability Engineer
Site Reliability Engineer

Spacex • Memphis (TN), Northern (KY)

Hybrid
USD 140,000 - 190,000
Site Reliability Engineer
Site Reliability Engineer

SpaceXAI • Southaven (MS)

On-site
USD 120,000 - 170,000
Software Engineer, Data Center - Memphis
Software Engineer, Data Center - Memphis

Xai • Southaven (MS)

On-site
USD 100,000 - 180,000
Software Engineer, Data Center - Memphis
Software Engineer, Data Center - Memphis

xAI • Memphis (TN)

On-site
USD 110,000 - 150,000
Site Reliability Engineer, Data Center - Memphis
Site Reliability Engineer, Data Center - Memphis

Xai • Southaven (MS)

On-site
USD 90,000 - 140,000
Site Reliability Engineer, Data Center - Memphis
Site Reliability Engineer, Data Center - Memphis

xAI • Memphis (TN)

On-site
USD 120,000 - 160,000
Network Operations Center Specialist - Memphis
Network Operations Center Specialist - Memphis

Pantera Capital • Southaven (MS)

On-site
USD 70,000 - 100,000
Software Engineer - Datacenter
Software Engineer - Datacenter

SpaceXAI • Memphis (TN)

On-site
USD 120,000 - 160,000
Site Reliability Engineer, Data Center - Memphis
Site Reliability Engineer, Data Center - Memphis

Pantera Capital • Southaven (MS)

On-site
USD 90,000 - 130,000