Staff Engineer, GPU Infrastructure & Automation

Matcha

Northern (KY)

Hybrid

USD 150,000 - 300,000

Full time

7 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Prime Intellect is building the open superintelligence stack. We are seeking an infrastructure engineer to own the machine lifecycle from discovery to secure reuse, enabling production-grade GPU servers.

You will automate provisioning, BIOS, BMC, drivers, and firmware; integrate with SLURM and Kubernetes; and collaborate with datacenter teams to keep systems healthy. Compensation ranges from $150k to $300k plus equity.

Qualifications

  • 3+ years of experience operating Linux servers or building bare-metal automation in production.
  • Hands-on experience with PXE/iPXE, DHCP, image provisioning, and out-of-band management such as Redfish or IPMI.
  • Strong software engineering and debugging skills in Python, Go, or a comparable language, plus Bash.
  • Experience designing reliable automation that handles partial failures, retries, and configuration drift.
  • Ability to own operational incidents and collaborate across hardware, networking, and platform teams.

Responsibilities

  • Build automated discovery, network boot, OS imaging, and configuration workflows for GPU servers.
  • Automate BIOS, BMC, NIC, GPU driver, and firmware configuration with staged rollouts and safe recovery paths.
  • Develop hardware inventory and lifecycle services that track machine identity, configuration, health, and readiness.
  • Create acceptance tests and burn-in workflows for GPUs, memory, storage, and interconnects before capacity enters production.
  • Integrate provisioning and health checks with SLURM, Kubernetes, and compute allocation systems.
  • Build observability, quarantine, repair, and re-provisioning workflows; partner with datacenter teams to resolve hardware failures.
  • Implement secure credential handling, tenant isolation, and data sanitization across the server lifecycle.

Skills

Linux servers
Python
Go
Bash
Automation design
Incident ownership

Tools

PXE/iPXE
Redfish/IPMI
Ansible
Terraform
Kubernetes
SLURM
MAAS
Ironic

Job description

Prime Intellect is building the open superintelligence stack. We are seeking an infrastructure engineer to own the machine lifecycle from discovery to secure reuse, enabling production-grade GPU servers.

You will automate provisioning, BIOS, BMC, drivers, and firmware; integrate with SLURM and Kubernetes; and collaborate with datacenter teams to keep systems healthy. Compensation ranges from $150k to $300k plus equity.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff Engineer Bare-Metal GPU Fleet Provisioning
Staff Engineer Bare-Metal GPU Fleet Provisioning

AI Chopping Block • San Francisco (CA), Northern (KY)

On-site
USD 150,000 - 300,000
Staff Engineer, Bare-Metal GPU Fleet Provisioning
Staff Engineer, Bare-Metal GPU Fleet Provisioning

Prime Intellect • San Francisco (CA)

On-site
USD 150,000 - 300,000
Member of Technical Staff - Bare Metal & Fleet Provisioning at Prime Intellect
Member of Technical Staff - Bare Metal & Fleet Provisioning at Prime Intellect

Matcha • Northern (KY)

Hybrid
USD 150,000 - 300,000
Member of Technical Staff - Bare Metal & Fleet Provisioning
Member of Technical Staff - Bare Metal & Fleet Provisioning

Prime Intellect • San Francisco (CA)

On-site
USD 150,000 - 300,000
GPU Cloud Infrastructure Engineer
GPU Cloud Infrastructure Engineer

Prime Intellect • San Francisco (CA)

On-site
USD 150,000 - 300,000
Member of Technical Staff - Bare Metal & Fleet Provisioning
Member of Technical Staff - Bare Metal & Fleet Provisioning

AI Chopping Block • San Francisco (CA), Northern (KY)

Hybrid
USD 150,000 - 300,000
Member of Technical Staff - Bare Metal & Fleet Provisioning
Member of Technical Staff - Bare Metal & Fleet Provisioning

Prime Intellect AI • San Francisco (CA)

On-site
USD 150,000 - 300,000
Member of Technical Staff - GPU Infrastructure
Member of Technical Staff - GPU Infrastructure

Prime Intellect • San Francisco (CA), Northern (KY)

Hybrid
USD 150,000 - 300,000
Equity incentives
Member of Technical Staff - GPU Infrastructure
Member of Technical Staff - GPU Infrastructure

Prime Intellect AI • San Francisco (CA)

On-site
USD 150,000 - 300,000
Member of Technical Staff - Datacenter Operations
Member of Technical Staff - Datacenter Operations

AI Chopping Block • San Francisco (CA), Northern (KY)

Hybrid
USD 150,000 - 300,000