Senior GPU Data Center Engineer

Prime Intellect AI

San Francisco (CA)

On-site

USD 150,000 - 300,000

Full time

6 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Prime Intellect AI in San Francisco seeks an operations-focused engineer to own the GPU cloud infrastructure. You will coordinate datacenter deployments, manage hardware, and drive incident response with our partners to ensure dependable production capacity.

You will maintain asset records, runbooks, and performance dashboards, while collaborating with facilities, networking, and software teams to scale high-density GPU deployments for frontier AI workloads.

Qualifications

  • 3+ years in datacenter operations, hardware infrastructure, or production systems operations.
  • Hands-on experience deploying and troubleshooting rack-mounted servers, networking equipment, and structured cabling.
  • Experience coordinating datacenter providers, remote hands, and hardware vendors through deployments and incidents.
  • Working knowledge of Linux diagnostics, BMC consoles, and server hardware health tools.
  • Strong operational judgment, documentation habits, and ownership of issues through resolution.

Responsibilities

  • Own the operational readiness of the GPU cloud infrastructure behind our platform.
  • Coordinate rack deployment, cabling, inventory, and acceptance testing with datacenter partners and engineering teams.
  • Maintain accurate asset records, rack layouts, power allocations, cabling documentation, and spare-parts inventories.
  • Lead hardware fault triage and coordinate remote hands, vendor escalations, component replacement, and RMA workflows.
  • Establish maintenance plans and change procedures that minimize customer disruption and protect equipment and data.
  • Track capacity readiness, hardware failure trends, repair times, and operational risks; automate reporting and workflows.
  • Partner with facility teams on power, cooling, environmental monitoring, and readiness for high-density GPU deployments.
  • Create runbooks and escalation procedures and support incident response across datacenter and infrastructure teams.

Skills

Linux diagnostics
BMC consoles
Server health tools
Rack deployment
Vendor coordination
Scripting basics

Tools

DGX/HGX systems
Fleet provisioning
Asset tracking software

Job description

Prime Intellect AI in San Francisco seeks an operations-focused engineer to own the GPU cloud infrastructure. You will coordinate datacenter deployments, manage hardware, and drive incident response with our partners to ensure dependable production capacity.

You will maintain asset records, runbooks, and performance dashboards, while collaborating with facilities, networking, and software teams to scale high-density GPU deployments for frontier AI workloads.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

GPU Cloud Infrastructure Engineer
GPU Cloud Infrastructure Engineer

Prime Intellect • San Francisco (CA)

On-site
USD 150,000 - 300,000
Senior GPU Cloud Operations Engineer
Senior GPU Cloud Operations Engineer

AI Chopping Block • San Francisco (CA), Northern (KY)

On-site
USD 150,000 - 300,000
Senior GPU Infrastructure Engineer — HPC & Clusters
Senior GPU Infrastructure Engineer — HPC & Clusters

Prime Intellect AI • San Francisco (CA)

On-site
USD 150,000 - 300,000
GPU Infrastructure Engineer | Datacenter Ops
GPU Infrastructure Engineer | Datacenter Ops

Prime Intellect • United States

Remote
USD 150,000 - 300,000
Member of Technical Staff - Datacenter Operations
Member of Technical Staff - Datacenter Operations

Prime Intellect • San Francisco (CA)

On-site
USD 150,000 - 300,000
Member of Technical Staff - Datacenter Operations
Member of Technical Staff - Datacenter Operations

AI Chopping Block • San Francisco (CA), Northern (KY)

Hybrid
USD 150,000 - 300,000
Member of Technical Staff - Datacenter Operations
Member of Technical Staff - Datacenter Operations

Prime Intellect AI • San Francisco (CA)

On-site
USD 150,000 - 300,000
Member of Technical Staff - Datacenter Operations
Member of Technical Staff - Datacenter Operations

Prime Intellect • United States

Remote
USD 150,000 - 300,000
Staff Datacenter Networking for GPU AI Infrastructure
Staff Datacenter Networking for GPU AI Infrastructure

Prime Intellect AI • San Francisco (CA)

On-site
USD 150,000 - 300,000
Staff Datacenter Networking Engineer for GPU Infra
Staff Datacenter Networking Engineer for GPU Infra

Prime Intellect • United States

On-site
USD 150,000 - 300,000