GPU Data Center Deployment Lead — Accelerate AI Clusters

SpaceXAI

Memphis (TN)

Hybrid

USD 150,000 - 230,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

SpaceXAI is seeking a Hardware Deployment Engineer Lead to own end‑to‑end bring‑up of GPU compute hardware across the world\'s largest AI training clusters. You will lead an in‑house team responsible for L11 integration, hardware bring‑up, and post‑L11 repair of GB300‑class systems across multiple data halls, driving deployment velocity and high availability.

The role is based in Memphis, TN, with a focus on setting deployment playbooks, ensuring SLA compliance with vendors, and scaling

Qualifications

  • 5+ years of hands‑on experience deploying, integrating, or repairing compute/server hardware at data center scale.
  • Direct experience with L11 (rack‑level) integration and bring‑up of GPU or accelerator‑based systems.
  • Demonstrated experience leading technician or engineering teams in a fast‑paced deployment, manufacturing, or data center environment.
  • Deep troubleshooting skills across servers, GPUs, NVLink/fabric interconnects, high‑speed networking, and liquid cooling systems.

Responsibilities

  • Lead, hire, and develop a dedicated hardware deployment team with full ownership of team structure and staffing.
  • Own L11 rack integration and compute hardware bring‑up across multiple data halls concurrently, from delivery dock to healthy production handoff.
  • Drive aggressive bring‑up timelines: achieve 95%+ node availability within days of rack delivery and 100% closure within one week per data hall.
  • Own post‑L11 hardware health: run systematic health pushes to sustain greater than 98% node availability prior to turnover to operations.
  • Internalize non‑RMA hardware repairs to maximize hardware recovery, minimize repair backlogs, and reduce dependence on OEM turnaround times.
  • Develop and enforce vendor SLAs for OEM and supplier responsibilities; prevent accumulation of unrepaired hardware and repair backlogs before turnover to operations.
  • Perform root cause analysis of hardware failures discovered during L11 and drive corrective actions with vendors and internal engineering teams.
  • Partner with site operations on hardware debugging and repair, and train site operations teams to support future data center deployments.
  • Build, document, and continuously improve deployment processes, tooling, and training so bring‑up capability scales across sites and future hardware generations.

Skills

Hands-on deployment
L11 integration
Team leadership
GPU troubleshooting
NVLink interconnects
Liquid cooling
Vendor management
Hardware bring-up

Tools

Burn-in tooling
Telemetry systems
Automation tooling

Job description

SpaceXAI is seeking a Hardware Deployment Engineer Lead to own end‑to‑end bring‑up of GPU compute hardware across the world\'s largest AI training clusters. You will lead an in‑house team responsible for L11 integration, hardware bring‑up, and post‑L11 repair of GB300‑class systems across multiple data halls, driving deployment velocity and high availability.

The role is based in Memphis, TN, with a focus on setting deployment playbooks, ensuring SLA compliance with vendors, and scaling

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Lead GPU Deployment Engineer - Data Center Scale
Lead GPU Deployment Engineer - Data Center Scale

Pantera Capital • Memphis (TN)

On-site
USD 140,000 - 210,000
Lead, Hardware Deployment Engineer
Lead, Hardware Deployment Engineer

Pantera Capital • Memphis (TN)

On-site
USD 140,000 - 210,000
Lead, Hardware Deployment Engineer
Lead, Hardware Deployment Engineer

SpaceXAI • Memphis (TN)

On-site
USD 150,000 - 230,000
Data Center Operations Lead for AI Compute
Data Center Operations Lead for AI Compute

Xai • Memphis (TN)

On-site
USD 110,000 - 150,000
Lead Engineer (Data Center)
Lead Engineer (Data Center)

Socket.dev • Memphis (TN)

On-site
USD 150,000 - 230,000
AI Data Center Ops Lead — GPU/HPC Infrastructure
AI Data Center Ops Lead — GPU/HPC Infrastructure

Nscale • Town of Norway (WI)

On-site
USD 120,000 - 170,000
Technical Account Manager & Solutions Architect, AI Compute & GPUs
Technical Account Manager & Solutions Architect, AI Compute & GPUs

xAI • Town of Texas (WI)

On-site
USD 140,000 - 210,000
Health insurance
Fertility benefits
Flexible vacation
+2
Head of Datacenter Compute Sourcing for Hyperscale AI
Head of Datacenter Compute Sourcing for Hyperscale AI

SpaceX • Austin (TX)

On-site
USD 250,000 - 350,000
Global AI Cluster Deployment Lead
Global AI Cluster Deployment Lead

Cerebras • Sunnyvale (CA)

On-site
USD 180,000 - 230,000
GPU Data Center Deployment Lead
GPU Data Center Deployment Lead

Dormont Manufacturing Co • Birmingham (AL)

On-site
USD 125,000 - 180,000
Health insurance
401(k) plan
Parental leave
+2