GPU Infrastructure Operations Lead

Primeintellect

San Francisco (CA)

On-site

USD 150,000 - 300,000

Full time

12 days ago
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Prime Intellect is building the open superintelligence stack. You’ll own the operational readiness of the physical infrastructure behind our GPU cloud, coordinating deployments, hardware maintenance, and incident response with datacenter partners, turning new capacity into dependable production infrastructure.

The role covers rack deployment, asset tracking, fault triage, and collaboration with facility teams to sustain high-density GPU deployments while enabling engineering and operations to

Qualifications

  • 3+ years in datacenter operations, hardware infrastructure, or production systems operations.
  • Hands-on experience deploying and troubleshooting rack-mounted servers, networking equipment, and structured cabling.
  • Experience coordinating datacenter providers, remote hands, and hardware vendors through deployments and incidents.
  • Working knowledge of Linux diagnostics, BMC consoles, and server hardware health tools.
  • Strong operational judgment, documentation habits, and ownership of issues through resolution.

Responsibilities

  • Coordinate rack deployment, cabling, inventory, and acceptance testing for new GPU capacity with datacenter partners and engineering teams.
  • Maintain accurate asset records, rack layouts, power allocations, cabling documentation, and spare-parts inventories.
  • Lead hardware fault triage and coordinate remote hands, vendor escalations, component replacement, and RMA workflows.
  • Establish maintenance plans and change procedures that minimize customer disruption and protect equipment and data.
  • Track capacity readiness, hardware failure trends, repair times, and operational risks; automate repetitive reporting and workflows.
  • Partner with facility teams on power, cooling, environmental monitoring, and readiness for high-density GPU deployments.
  • Create runbooks and escalation procedures and support incident response across datacenter and infrastructure teams.

Skills

Datacenter ops
Rack deployment
Linux diagnostics
BMC consoles
Asset tracking
Vendor coordination
Scripting
Hardware fault triage
Power & cooling management

Tools

BMC consoles

Job description

Prime Intellect is building the open superintelligence stack. You’ll own the operational readiness of the physical infrastructure behind our GPU cloud, coordinating deployments, hardware maintenance, and incident response with datacenter partners, turning new capacity into dependable production infrastructure.

The role covers rack deployment, asset tracking, fault triage, and collaboration with facility teams to sustain high-density GPU deployments while enabling engineering and operations to

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

GPU Infrastructure Engineer | Datacenter Ops
GPU Infrastructure Engineer | Datacenter Ops

Prime Intellect • United States

Remote
USD 150,000 - 300,000
GPU Cloud Infrastructure Engineer
GPU Cloud Infrastructure Engineer

Prime Intellect • San Francisco (CA)

On-site
USD 150,000 - 300,000
Senior GPU Cloud Operations Engineer
Senior GPU Cloud Operations Engineer

AI Chopping Block • San Francisco (CA), Northern (KY)

On-site
USD 150,000 - 300,000
Senior GPU Data Center Engineer
Senior GPU Data Center Engineer

Prime Intellect AI • San Francisco (CA)

On-site
USD 150,000 - 300,000
Senior Bare-Metal GPU Infrastructure Engineer
Senior Bare-Metal GPU Infrastructure Engineer

Primeintellect • San Francisco (CA)

On-site
USD 150,000 - 300,000
Senior GPU Infrastructure Architect
Senior GPU Infrastructure Architect

Primeintellect • San Francisco (CA)

On-site
USD 150,000 - 300,000
Staff Engineer, GPU Infrastructure & Automation
Staff Engineer, GPU Infrastructure & Automation

Matcha • Northern (KY)

Hybrid
USD 150,000 - 300,000
Member of Technical Staff - Datacenter Operations
Member of Technical Staff - Datacenter Operations

Prime Intellect • United States

Remote
USD 150,000 - 300,000
Member of Technical Staff - Datacenter Operations
Member of Technical Staff - Datacenter Operations

Prime Intellect • San Francisco (CA)

On-site
USD 150,000 - 300,000
Member of Technical Staff - Datacenter Operations
Member of Technical Staff - Datacenter Operations

Primeintellect • San Francisco (CA)

On-site
USD 150,000 - 300,000