Staff Engineer, Bare-Metal GPU Fleet Provisioning

Prime Intellect

San Francisco (CA)

On-site

USD 150,000 - 300,000

Full time

4 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Prime Intellect in San Francisco builds the open frontier AI infrastructure and seeks an Infrastructure Engineer to own the machine lifecycle, turning bare-metal GPU servers into production-ready compute. You will manage discovery, provisioning, validation, upgrades, repair, and secure reuse.

You’ll implement automated discovery, imaging, BIOS/firmware configuration, and health checks, integrating with SLURM, Kubernetes, and compute allocation systems to scale deployments.

Qualifications

  • Experience operating Linux servers in production environments.
  • Hands-on with PXE/iPXE, DHCP, image provisioning and out-of-band management such as Redfish/IPMI.
  • Strong software engineering skills in Python/Go and Bash.
  • Ability to design automation that handles partial failures, retries and drift.

Responsibilities

  • Build automated discovery, network boot, OS imaging, and configuration workflows for GPU servers.
  • Automate BIOS, BMC, NIC, GPU driver, and firmware configurations with safe rollouts.
  • Develop hardware inventory and lifecycle services tracking identity, config, health, readiness.
  • Create burn-in tests for GPUs, memory, storage, and interconnects before production.
  • Integrate provisioning and health checks with SLURM, Kubernetes, and compute allocation systems.
  • Build observability, repair, and re-provisioning workflows; collaborate with data center teams.
  • Implement secure credential handling and tenant isolation across server lifecycle.

Skills

Linux servers
Bare-metal automation
PXE/iPXE
DHCP
Image provisioning
Redfish/IPMI
Python
Go
Bash
Incident ownership

Tools

Ansible
Terraform
Kubernetes
SLURM

Job description

Prime Intellect in San Francisco builds the open frontier AI infrastructure and seeks an Infrastructure Engineer to own the machine lifecycle, turning bare-metal GPU servers into production-ready compute. You will manage discovery, provisioning, validation, upgrades, repair, and secure reuse.

You’ll implement automated discovery, imaging, BIOS/firmware configuration, and health checks, integrating with SLURM, Kubernetes, and compute allocation systems to scale deployments.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

GPU Cloud Infrastructure Engineer
GPU Cloud Infrastructure Engineer

Prime Intellect • San Francisco (CA)

On-site
USD 150,000 - 300,000
Member of Technical Staff - Bare Metal & Fleet Provisioning
Member of Technical Staff - Bare Metal & Fleet Provisioning

Prime Intellect • San Francisco (CA)

On-site
USD 150,000 - 300,000
Member of Technical Staff - Bare Metal & Fleet Provisioning
Member of Technical Staff - Bare Metal & Fleet Provisioning

Prime-Intellect • San Francisco (CA)

On-site
USD 150,000 - 300,000
Software Engineer, Compute Foundations
Software Engineer, Compute Foundations

Linuxcareers • San Francisco (CA), Northern (KY)

Hybrid
USD 210,000 - 270,000
Senior Infrastructure Engineer
Senior Infrastructure Engineer

Harrison Clarke • United States

On-site
USD 100,000 - 140,000
GPU Compute Engineer — Fleet Reliability & Automation
GPU Compute Engineer — Fleet Reliability & Automation

Fluidstack • San Francisco (CA)

On-site
USD 175,000 - 300,000
Health, dental, and vision insurance
Retirement or pension plan
Generous PTO policy
Autonomous AI Infrastructure Engineer: GPU Fleet Mastery
Autonomous AI Infrastructure Engineer: GPU Fleet Mastery

Together • San Francisco (CA)

On-site
USD 190,000 - 270,000
Health insurance
Startup equity
Competitive benefits
Senior Cloud Infrastructure Engineer – GPU & DPU
Senior Cloud Infrastructure Engineer – GPU & DPU

Lambda • United States

Remote
USD 180,000 - 260,000
AI Infra Systems Engineer: GPU & Accelerator Servers
AI Infra Systems Engineer: GPU & Accelerator Servers

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 129,000 - 175,000
Health insurance
401(k) matching
Paid time off
+1
AI Infra Engineer: GPU Fleet Automation
AI Infra Engineer: GPU Fleet Automation

Fal.ai Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 250,000
Relocation assistance
Health, dental, and vision insurance (
Regular team events and offsites