Senior GPU Cloud Operations Engineer

AI Chopping Block

San Francisco, Northern (CA, KY)

On-site

USD 150,000 - 300,000

Full time

6 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Prime Intellect is building the open frontier AI stack and operates a GPU cloud infrastructure that powers frontier models and research. The role focuses on ensuring operational readiness, coordinating datacenter deployments, and maintaining high-availability infrastructure for production workloads.

You will work with a hands-on team to minimize disruption, automate reporting, and drive reliability across large-scale GPU deployments.

Qualifications

  • 3+ years in datacenter operations, hardware infrastructure, or production systems operations.
  • Hands-on experience deploying and troubleshooting rack-mounted servers, networking equipment, and structured cabling.
  • Experience coordinating datacenter providers, remote hands, and hardware vendors through deployments and incidents.
  • Working knowledge of Linux diagnostics, BMC consoles, and server hardware health tools.
  • Strong operational judgment, documentation habits, and ownership of issues through resolution.

Responsibilities

  • Coordinate rack deployment, cabling, inventory, and acceptance testing for new GPU capacity with datacenter partners and engineering teams
  • Maintain accurate asset records, rack layouts, power allocations, cabling documentation, and spare-parts inventories
  • Lead hardware fault triage and coordinate remote hands, vendor escalations, component replacement, and RMA workflows
  • Establish maintenance plans and change procedures that minimize customer disruption and protect equipment and data
  • Track capacity readiness, hardware failure trends, repair times, and operational risks; automate repetitive reporting and workflows
  • Partner with facility teams on power, cooling, environmental monitoring, and readiness for high-density GPU deployments
  • Create runbooks and escalation procedures and support incident response across datacenter and infrastructure teams

Skills

Datacenter operations
Hardware infrastructure
Linux diagnostics
Asset tracking
Vendor coordination

Tools

BMC consoles
Server health tools
Scripting (bash/python)

Job description

Prime Intellect is building the open frontier AI stack and operates a GPU cloud infrastructure that powers frontier models and research. The role focuses on ensuring operational readiness, coordinating datacenter deployments, and maintaining high-availability infrastructure for production workloads.

You will work with a hands-on team to minimize disruption, automate reporting, and drive reliability across large-scale GPU deployments.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

GPU Cloud Infrastructure Engineer
GPU Cloud Infrastructure Engineer

Prime Intellect • San Francisco (CA)

On-site
USD 150,000 - 300,000
GPU Infrastructure Engineer | Datacenter Ops
GPU Infrastructure Engineer | Datacenter Ops

Prime Intellect • United States

Remote
USD 150,000 - 300,000
Staff Datacenter Networking Engineer for GPU Infra
Staff Datacenter Networking Engineer for GPU Infra

Prime Intellect • United States

On-site
USD 150,000 - 300,000
Staff Datacenter Networking Engineer: Frontier AI GPU
Staff Datacenter Networking Engineer: Frontier AI GPU

Prime Intellect • San Francisco (CA)

On-site
USD 150,000 - 300,000
Member of Technical Staff - Datacenter Operations
Member of Technical Staff - Datacenter Operations

Prime Intellect AI • San Francisco (CA)

On-site
USD 150,000 - 300,000
Member of Technical Staff - Datacenter Operations
Member of Technical Staff - Datacenter Operations

AI Chopping Block • San Francisco (CA), Northern (KY)

Hybrid
USD 150,000 - 300,000
Member of Technical Staff - Datacenter Operations
Member of Technical Staff - Datacenter Operations

Prime Intellect • San Francisco (CA)

On-site
USD 150,000 - 300,000
Member of Technical Staff - Datacenter Operations
Member of Technical Staff - Datacenter Operations

Prime Intellect • United States

Remote
USD 150,000 - 300,000
Staff Engineer Bare-Metal GPU Fleet Provisioning
Staff Engineer Bare-Metal GPU Fleet Provisioning

AI Chopping Block • San Francisco (CA), Northern (KY)

On-site
USD 150,000 - 300,000
Senior GPU Cloud Infrastructure Engineer
Senior GPU Cloud Infrastructure Engineer

Hyperbolic • San Francisco (CA)

On-site
USD 180,000 - 240,000