Staff Engineer - Bare-Metal & GPU Fleet Provisioning

Prime Intellect

United States

On-site

USD 150,000 - 300,000

Full time

2 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Prime Intellect in the United States is seeking an experienced infrastructure engineer to turn bare-metal GPU servers into reliable, production-ready compute. You will own the machine lifecycle from discovery and provisioning through validation, upgrades, repair, and secure reuse as our fleet grows.

You will build automated workflows for GPU servers, automate BIOS, NIC, and firmware, develop inventory and lifecycle services, and integrate with SLURM, Kubernetes, and compute allocation systems to

Qualifications

  • 3+ years of experience operating Linux servers or building bare-metal infrastructure automation in production.
  • Hands-on experience with PXE/iPXE, DHCP, image provisioning, and out-of-band management such as Redfish or IPMI.
  • Strong software engineering and debugging skills in Python, Go, or a comparable language, plus Bash.
  • Experience designing reliable automation that handles partial failures, retries, and configuration drift.
  • Ability to own operational incidents and collaborate across hardware, networking, and platform teams.

Responsibilities

  • Build automated discovery, network boot, OS imaging, and configuration workflows for GPU servers.
  • Automate BIOS, BMC, NIC, GPU driver, and firmware configuration with staged rollouts and safe recovery paths.
  • Develop hardware inventory and lifecycle services that track machine identity, configuration, health, and readiness.
  • Create acceptance tests and burn-in workflows for GPUs, memory, storage, and interconnects before capacity enters production.
  • Integrate provisioning and health checks with SLURM, Kubernetes, and compute allocation systems.
  • Build observability, quarantine, repair, and re-provisioning workflows; partner with datacenter teams to resolve hardware failures.
  • Implement secure credential handling, tenant isolation, and data sanitization across the server lifecycle.

Skills

Linux servers
Infrastructure automation
Python
Go
Bash
Partial failure design
Incident ownership

Tools

Ansible
Terraform
MAAS
Ironic
Tinkerbell

Job description

Prime Intellect in the United States is seeking an experienced infrastructure engineer to turn bare-metal GPU servers into reliable, production-ready compute. You will own the machine lifecycle from discovery and provisioning through validation, upgrades, repair, and secure reuse as our fleet grows.

You will build automated workflows for GPU servers, automate BIOS, NIC, and firmware, develop inventory and lifecycle services, and integrate with SLURM, Kubernetes, and compute allocation systems to

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff Engineer, Bare-Metal GPU Fleet Provisioning
Staff Engineer, Bare-Metal GPU Fleet Provisioning

Prime Intellect • San Francisco (CA)

On-site
USD 150,000 - 300,000
Staff Engineer Bare-Metal GPU Fleet Provisioning
Staff Engineer Bare-Metal GPU Fleet Provisioning

AI Chopping Block • San Francisco (CA), Northern (KY)

On-site
USD 150,000 - 300,000
GPU Infrastructure Engineer | Datacenter Ops
GPU Infrastructure Engineer | Datacenter Ops

Prime Intellect • United States

Remote
USD 150,000 - 300,000
Staff Engineer, GPU Infrastructure & Automation
Staff Engineer, GPU Infrastructure & Automation

Matcha • Northern (KY)

Hybrid
USD 150,000 - 300,000
GPU Cloud Infrastructure Engineer
GPU Cloud Infrastructure Engineer

Prime Intellect • San Francisco (CA)

On-site
USD 150,000 - 300,000
Bare-Metal GPU Support Engineer
Bare-Metal GPU Support Engineer

CoreWeave • Livingston (NJ)

On-site
USD 83,000 - 145,000
Medical, dental, and vision insurance
401(k) with company match
Equity awards
+2
Staff Network Engineer — GPU Data Center & HPC Networking
Staff Network Engineer — GPU Data Center & HPC Networking

Matcha • Northern (KY)

Hybrid
USD 150,000 - 300,000
GPU Bare-Metal & DPU Engineer — Automation & Ops
GPU Bare-Metal & DPU Engineer — Automation & Ops

Bitdeer (NASDAQ: BTDR) • San Jose (CA)

On-site
USD 170,000 - 250,000
Member of Technical Staff - Bare Metal & Fleet Provisioning
Member of Technical Staff - Bare Metal & Fleet Provisioning

Prime Intellect • San Francisco (CA)

On-site
USD 150,000 - 300,000
Senior Cloud Infrastructure Engineer – GPU & DPU
Senior Cloud Infrastructure Engineer – GPU & DPU

Lambda • United States

Remote
USD 180,000 - 260,000