Infrastructure Engineer

Acceler8 Talent

San Francisco, Northern (CA, KY)

Hybrid

USD 150,000 - 390,000

Full time

8 days ago
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Equity

Job summary

Acceler8 Talent is building a cloud platform that runs inference workloads across GPUs, CPUs, and accelerators. You’ll design how new hardware is brought online and how compute fleets are provisioned and operated.

This role focuses on production-ready infrastructure with observability, reliable scheduling, and close collaboration with runtime, compiler, distributed-systems, networking, and hardware engineers.

Qualifications

  • Experience in building and operating large-scale infrastructure
  • Strong Linux systems knowledge and production debugging experience
  • Experience operating Kubernetes, Slurm, Nomad, or similar systems
  • Infrastructure automation experience using Python, Go, Terraform, or Ansible
  • Experience with GPU/accelerator infrastructure, including drivers, firmware, CUDA, or ROCm
  • Track record of observable, recoverable, and reliable production systems

Responsibilities

  • Build the cluster infrastructure behind the AI inference platform
  • Determine how new hardware is brought online and how compute fleets are provisioned
  • Improve cluster scheduling, resource utilization, isolation, and capacity management
  • Develop observability systems for faster debugging and recovery
  • Make accelerators production-ready across drivers, firmware, networking, and orchestration
  • Collaborate with runtime, compiler, distributed-systems, networking, and hardware engineers

Skills

Infrastructure
Platform engineering
SRE
HPC
Linux systems knowledge
Production debugging
Observability
Cluster engineering

Tools

Kubernetes
Slurm
Nomad
Python
Go
Terraform
Ansible
CUDA
ROCm

Job description

Member of Technical Staff, Infrastructure

On-site | San Francisco, CA | 5 days per week

$150k to $390k base + equity

I’m working with a well-funded AI infrastructure startup (Series A) building a cloud platform that runs inference workloads across GPUs, CPUs, and emerging accelerator architectures.

The team is solving a difficult infrastructure problem: making new different hardware accelerators usable through one reliable platform, without requiring customers to redesign their software stack for every accelerator.

This role will build the cluster infrastructure behind that platform. You’ll determine how new hardware is brought online, how compute fleets are provisioned and operated, and how production inference systems remain reliable as they scale.

You’ll work on problems such as:

  • Deploying production clusters across different accelerator architectures
  • Automating bare-metal provisioning, validation, upgrades, and fleet lifecycle management
  • Improving cluster scheduling, resource utilization, isolation, and capacity management
  • Building observability systems for faster debugging, incident response, and recovery
  • Making new accelerators production-ready across drivers, firmware, networking, and orchestration
  • Partnering with runtime, compiler, distributed-systems, networking, and hardware engineers.

Looking for engineers who have:

  • Experience in infrastructure, platform engineering, cluster engineering, SRE, or HPC
  • Strong Linux systems knowledge and production debugging experience
  • Experience operating Kubernetes, Slurm, Nomad, or similar orchestration systems
  • Infrastructure automation experience using Python, Go, Terraform, or Ansible
  • Experience with GPU or accelerator infrastructure, including drivers, firmware, CUDA, or ROCm
  • A track record of building observable, recoverable, and reliable production systems

This is an opportunity to join a small, highly technical team and build production infrastructure across multiple generations and types of AI hardware. The strongest candidates will be able to explain what they personally built, operated, measured, and debugged at scale.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Member of Technical Staff - Infrastructure
Member of Technical Staff - Infrastructure

Gimlet Labs • San Francisco (CA)

On-site
USD 120,000 - 160,000
Head of AI Data Center Infrastructure Platforms and Software
Head of AI Data Center Infrastructure Platforms and Software

Summit Group Solutions, LLC • United States

On-site
USD 150,000 - 350,000
Infrastructure / Cluster Engineer
Infrastructure / Cluster Engineer

Linuxcareers • San Francisco (CA)

On-site
USD 120,000 - 160,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

The Recruiting Guy • Arlington (VA)

On-site
USD 175,000 - 250,000
Software Engineer, Compute Infrastructure
Software Engineer, Compute Infrastructure

OpenAI • Los Angeles (CA)

On-site
USD 230,000 - 405,000
Equity
Flexible work environment
Health benefits
Senior Cloud Infrastructure Engineer
Senior Cloud Infrastructure Engineer

The Recruiting Guy • New York (NY)

On-site
USD 175,000 - 250,000
Senior Solutions Engineer, AI Infrastructure
Senior Solutions Engineer, AI Infrastructure

VAST Data • New York (NY)

On-site
USD 150,000 - 200,000
Member of Technical Staff - GPU Infrastructure
Member of Technical Staff - GPU Infrastructure

Prime Intellect • United States

On-site
USD 120,000 - 150,000
Member of Technical Staff - GPU Infrastructure
Member of Technical Staff - GPU Infrastructure

Prime Intellect AI • San Francisco (CA)

On-site
USD 150,000 - 300,000
Cluster Design
Cluster Design

Blue Signal Search • San Francisco (CA)

On-site
USD 150,000 - 230,000