Infrastructure Engineer

Acceler8 Talent

San Francisco (CA)

On-site

USD 150,000 - 390,000

Full time

2 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Acceler8 Talent in San Francisco is seeking a Member of Technical Staff, Infrastructure to build and operate the cluster infrastructure behind our AI platform. You will determine how new accelerator hardware is brought online, how compute fleets are provisioned, and how production inference systems remain reliable at scale.

You will deploy production clusters across accelerator architectures, automate provisioning and fleet lifecycle, improve scheduling and resource utilization, and build

Qualifications

  • Experience in infrastructure, platform engineering, cluster engineering, SRE, or HPC.
  • Strong Linux systems knowledge and production debugging experience.
  • Experience operating Kubernetes, Slurm, Nomad, or similar orchestration systems.
  • Infrastructure automation experience using Python, Go, Terraform, or Ansible.
  • Experience with GPU or accelerator infrastructure, including drivers, firmware, CUDA, or ROCm.
  • A track record of building observable, recoverable, and reliable production systems.

Responsibilities

  • Deploy production clusters across different accelerator architectures.
  • Automate bare-metal provisioning, validation, upgrades, and fleet lifecycle management.
  • Improve cluster scheduling, resource utilization, isolation, and capacity management.
  • Build observability systems for faster debugging, incident response, and recovery.
  • Make new accelerators production-ready across drivers, firmware, networking, and orchestration.
  • Partner with runtime, compiler, distributed-systems, networking, and hardware engineers.

Skills

Linux systems
Production debugging
Observability
Cloud platforms

Tools

Kubernetes
Slurm
Nomad
Python
Go
Terraform
Ansible
CUDA
GPU drivers

Job description

Member of Technical Staff, Infrastructure
$150k to $390k base + equity

I’m working with a well-funded AI infrastructure startup (Series A) building a cloud platform that runs inference workloads across GPUs, CPUs, and emerging accelerator architectures.

The team is solving a difficult infrastructure problem: making new different hardware accelerators usable through one reliable platform, without requiring customers to redesign their software stack for every accelerator.

This role will build the cluster infrastructure behind that platform. You’ll determine how new hardware is brought online, how compute fleets are provisioned and operated, and how production inference systems remain reliable as they scale.

You’ll work on problems such as:
  • Deploying production clusters across different accelerator architectures
  • Automating bare-metal provisioning, validation, upgrades, and fleet lifecycle management
  • Improving cluster scheduling, resource utilization, isolation, and capacity management
  • Building observability systems for faster debugging, incident response, and recovery
  • Making new accelerators production-ready across drivers, firmware, networking, and orchestration
  • Partnering with runtime, compiler, distributed-systems, networking, and hardware engineers.
Looking for engineers who have:
  • Experience in infrastructure, platform engineering, cluster engineering, SRE, or HPC
  • Strong Linux systems knowledge and production debugging experience
  • Experience operating Kubernetes, Slurm, Nomad, or similar orchestration systems
  • Infrastructure automation experience using Python, Go, Terraform, or Ansible
  • Experience with GPU or accelerator infrastructure, including drivers, firmware, CUDA, or ROCm
  • A track record of building observable, recoverable, and reliable production systems

This is an opportunity to join a small, highly technical team and build production infrastructure across multiple generations and types of AI hardware. The strongest candidates will be able to explain what they personally built, operated, measured, and debugged at scale.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Member of Technical Staff - Infrastructure
Member of Technical Staff - Infrastructure

Gimlet Labs • San Francisco (CA)

On-site
USD 120,000 - 160,000
Software Engineer - AI Infrastructure
Software Engineer - AI Infrastructure

Hamilton Barnes Associates Limited • San Francisco (CA)

On-site
USD 300,000 - 500,000
Early-stage equity
Founding engineer role
Equity package
Head of AI Data Center Infrastructure Platforms and Software
Head of AI Data Center Infrastructure Platforms and Software

Summit Group Solutions, LLC • United States

On-site
USD 150,000 - 350,000
Member of Technical Staff - AI Infrastructure
Member of Technical Staff - AI Infrastructure

Hamilton Barnes Associates Limited • San Francisco (CA)

On-site
USD 213,000 - 288,000
Early-stage equity
Direct access to leadership
Senior Site Reliability Engineer (SRE) - AI Inftastructure
Senior Site Reliability Engineer (SRE) - AI Inftastructure

Hamilton Barnes Associates Limited • San Francisco (CA)

On-site
USD 270,000 - 330,000
Equity
Infrastructure / Cluster Engineer
Infrastructure / Cluster Engineer

Linuxcareers • San Francisco (CA)

On-site
USD 120,000 - 160,000
Infrastructure Product Engineer - AI Infrastructure
Infrastructure Product Engineer - AI Infrastructure

Hamilton Barnes Associates Limited • San Francisco (CA)

On-site
USD 300,000 - 350,000
Equity
Cluster Design
Cluster Design

Blue Signal Search • San Francisco (CA)

On-site
USD 150,000 - 230,000
Member of Technical Staff - GPU Infrastructure
Member of Technical Staff - GPU Infrastructure

Prime Intellect • United States

On-site
USD 120,000 - 150,000
Member of Technical Staff - GPU Infrastructure
Member of Technical Staff - GPU Infrastructure

Prime Intellect • San Francisco (CA)

On-site
USD 150,000 - 300,000