Infrastructure Engineer, AI Hardware Platform & Equity

Acceler8 Talent

San Francisco (CA)

On-site

USD 150,000 - 390,000

Full time

3 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Acceler8 Talent in San Francisco is seeking a Member of Technical Staff, Infrastructure to build and operate the cluster infrastructure behind our AI platform. You will determine how new accelerator hardware is brought online, how compute fleets are provisioned, and how production inference systems remain reliable at scale.

You will deploy production clusters across accelerator architectures, automate provisioning and fleet lifecycle, improve scheduling and resource utilization, and build

Qualifications

  • Experience in infrastructure, platform engineering, cluster engineering, SRE, or HPC.
  • Strong Linux systems knowledge and production debugging experience.
  • Experience operating Kubernetes, Slurm, Nomad, or similar orchestration systems.
  • Infrastructure automation experience using Python, Go, Terraform, or Ansible.
  • Experience with GPU or accelerator infrastructure, including drivers, firmware, CUDA, or ROCm.
  • A track record of building observable, recoverable, and reliable production systems.

Responsibilities

  • Deploy production clusters across different accelerator architectures.
  • Automate bare-metal provisioning, validation, upgrades, and fleet lifecycle management.
  • Improve cluster scheduling, resource utilization, isolation, and capacity management.
  • Build observability systems for faster debugging, incident response, and recovery.
  • Make new accelerators production-ready across drivers, firmware, networking, and orchestration.
  • Partner with runtime, compiler, distributed-systems, networking, and hardware engineers.

Skills

Linux systems
Production debugging
Observability
Cloud platforms

Tools

Kubernetes
Slurm
Nomad
Python
Go
Terraform
Ansible
CUDA
GPU drivers

Job description

Acceler8 Talent in San Francisco is seeking a Member of Technical Staff, Infrastructure to build and operate the cluster infrastructure behind our AI platform. You will determine how new accelerator hardware is brought online, how compute fleets are provisioned, and how production inference systems remain reliable at scale.

You will deploy production clusters across accelerator architectures, automate provisioning and fleet lifecycle, improve scheduling and resource utilization, and build

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Infrastructure Engineer
Infrastructure Engineer

Acceler8 Talent • San Francisco (CA)

On-site
USD 150,000 - 390,000
Senior Staff AI Infra Engineer — Lead Scalable Compute
Senior Staff AI Infra Engineer — Lead Scalable Compute

Anthropic • San Francisco (CA)

Hybrid
USD 405,000 - 485,000
AI Infrastructure Engineer for Hardware & Tooling
AI Infrastructure Engineer for Hardware & Tooling

Oho Group • San Francisco (CA)

On-site
USD 150,000 - 190,000
Heterogeneous AI Infra & Cluster Engineer
Heterogeneous AI Infra & Cluster Engineer

Gimlet Labs • San Francisco (CA)

On-site
USD 120,000 - 160,000
AI Platform Engineer — Scale Cloud Infra, Equity
AI Platform Engineer — Scale Cloud Infra, Equity

Harrison Clarke • San Francisco (CA)

On-site
USD 150,000 - 210,000
Staff Engineer: AI Infrastructure & Scale Leader
Staff Engineer: AI Infrastructure & Scale Leader

Recruiting From Scratch • San Francisco (CA)

On-site
USD 250,000 - 300,000
Competitive equity
Ownership in infra
Direct collaboration with founders
+1
Senior AI Infrastructure Engineer — Scale GPU Clusters
Senior AI Infrastructure Engineer — Scale GPU Clusters

AI Breaking Wire • San Francisco (CA)

On-site
USD 280,000 - 400,000
Equity
Medical, dental, and vision benefits
Unlimited PTO
+2
Staff Software Engineer AI Inference Infra Orchestrator
Staff Software Engineer AI Inference Infra Orchestrator

Together AI • San Francisco (CA)

On-site
USD 240,000 - 280,000
Equity
Health insurance
Competitive benefits
Senior AI Infrastructure Engineer
Senior AI Infrastructure Engineer

AI Breaking Wire • Menlo Park (CA), Northern (KY)

Hybrid
USD 200,000 - 350,000
RSUs
Health benefits
Parental leave
+1
Senior Platform Engineer, Inference & GPU Compute Infra
Senior Platform Engineer, Inference & GPU Compute Infra

Together • San Francisco (CA)

On-site
USD 240,000 - 280,000
Startup equity
Health insurance