Remote GPU Cluster Engineer & Automation Lead

AI Chopping Block

Northern (KY)

Hybrid

USD 150,000 - 230,000

Full time

9 days ago
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

ElevenLabs seeks an experienced HPC Infrastructure Engineer to own our GPU cluster infrastructure from provisioning to performance tuning. You will build automation to keep the fleet healthy with minimal human intervention and manage OS images, NVIDIA drivers, CUDA, container runtimes, and NCCL.

You’ll run and optimize Slurm-like job scheduling, design fast storage, and diagnose bottlenecks across hardware and software.

Qualifications

  • Experience running large-scale Linux server or GPU environments in production.
  • Deep NVIDIA stack knowledge (drivers, CUDA, NCCL, DCGM) or equivalent systems experience.
  • Comfort with bare-metal environments, server hardware and high-speed networking.
  • Solid automation skills in Python and/or Bash; IaC tooling like Ansible or Terraform.
  • Ability to analyze metrics/logs to diagnose issues and improve performance.
  • Willingness to own end-to-end scope including datacenter tasks.

Responsibilities

  • Operate and optimize GPU clusters from provisioning to capacity planning.
  • Develop automation to maintain fleet health with minimal human intervention.
  • Manage OS images, NVIDIA drivers, CUDA, container runtimes, and NCCL.
  • Run and tune scheduling for researchers to maximize compute efficiency.
  • Build high-performance storage for datasets and checkpoints.
  • Identify and fix performance bottlenecks across hardware and software.
  • Evaluate rented GPU capacity and ensure provider SLAs are met.
  • Perform hands-on hardware work when needed (racking, cabling, diagnostics).
  • Maintain cluster security through access control and network isolation.

Skills

Linux server administration
NVIDIA stack expertise
Automation with Python/Bash
Metrics & logs analysis (PromQL)
End-to-end ownership
Datacenter hardware handling

Tools

Ansible
Terraform
PXE provisioning
BMC/IPMI/Redfish automation

Job description

ElevenLabs seeks an experienced HPC Infrastructure Engineer to own our GPU cluster infrastructure from provisioning to performance tuning. You will build automation to keep the fleet healthy with minimal human intervention and manage OS images, NVIDIA drivers, CUDA, container runtimes, and NCCL.

You’ll run and optimize Slurm-like job scheduling, design fast storage, and diagnose bottlenecks across hardware and software.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Remote HPC Infra Engineer — GPU Clusters
Remote HPC Infra Engineer — GPU Clusters

ElevenLabs • Northern (KY)

Hybrid
USD 140,000 - 210,000
HPC Infrastructure Engineer - GPU Clusters
HPC Infrastructure Engineer - GPU Clusters

AI Chopping Block • Northern (KY)

Hybrid
USD 150,000 - 230,000
HPC Infrastructure Engineer - GPU Clusters
HPC Infrastructure Engineer - GPU Clusters

ElevenLabs • Northern (KY)

Hybrid
USD 140,000 - 210,000
Senior HPC & GPU Cluster Architect — Scale & Automate
Senior HPC & GPU Cluster Architect — Scale & Automate

San Francisco Compute Company • San Francisco (CA)

Hybrid
USD 120,000 - 160,000
Generous equity grant
Competitive salary
Visa sponsorship
+6
Senior HPC Support Engineer - GPU Cloud Infra
Senior HPC Support Engineer - GPU Cloud Infra

Neura Market • United States

On-site
USD 150,000 - 190,000
Wellness stipend
Commuter stipend
401k plan with 2% company match (USA)
+1
Staff Engineer, Distributed GPU Clusters
Staff Engineer, Distributed GPU Clusters

Kindredventures • San Francisco (CA)

On-site
USD 140,000 - 230,000
Staff Engineer, Distributed GPU Clusters & Infra
Staff Engineer, Distributed GPU Clusters & Infra

Causal Labs • San Francisco (CA)

On-site
USD 180,000 - 240,000
Senior GPU Cluster Engineer for AI Infrastructure
Senior GPU Cluster Engineer for AI Infrastructure

Sciforium • San Francisco (CA)

On-site
USD 150,000 - 220,000
Medical, dental, and vision insurance
401k plan
Daily lunch, snacks, and beverages
+2
AI Kernel / Cluster Engineer
AI Kernel / Cluster Engineer

Blue Signal Search • Santa Clara (CA)

On-site
USD 150,000 - 210,000
Senior GPU Compute Cluster Engineer
Senior GPU Compute Cluster Engineer

Inferact Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 200,000 - 400,000
Health benefits
Dental benefits
Vision benefits
+1