Remote HPC Infra Engineer — GPU Clusters

ElevenLabs

Northern (KY)

Hybrid

USD 140,000 - 210,000

Full time

8 days ago
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

ElevenLabs is hiring for a remote HPC-Infrastructure Engineer specialized in GPU clusters. You’ll own GPU fleets, automate provisioning, scheduling, and monitoring, and ensure fast, reliable research compute across NVIDIA stacks and data centers.

Responsibilities include OS images, drivers, CUDA, NCCL, and high-speed networking, plus secure operations. We value hands-on engineers who thrive on end-to-end scope, enjoy building automation in Python/Bash, and can work with Slurm, Terraform, and

Qualifications

  • Experience running large-scale Linux or GPU environments in production.
  • Deep NVIDIA stack knowledge or ability to learn hardware stacks quickly.
  • Comfort with bare-metal environments and high-speed networking.
  • Proficient in Python/Bash automation; IaC tools like Ansible or Terraform.
  • Ability to analyze metrics and logs (PromQL) to diagnose issues.
  • Thrives with end-to-end ownership and datacenter involvement.

Responsibilities

  • Operate and improve the GPU fleet end to end: provisioning, scheduling, monitoring, upgrades.
  • Build automation to keep fleet healthy with automated checks and remediation.
  • Own the stack beneath training code: OS images, drivers, CUDA, containers.
  • Run and tune job scheduling (Slurm or similar) for fair, fast compute.
  • Build/maintain high-performance storage for datasets and checkpoints.
  • Diagnose performance issues and fix root causes, not just symptoms.
  • Evaluate rented GPU capacity and verify SLAs with providers.
  • Hands-on hardware tasks: racking, cabling, and datacenter coordination.
  • Ensure clusters are secure by default: access control and network isolation.

Skills

Large-scale Linux
NVIDIA stack mastery
Bare-metal & high-speed networking
Python/Bash automation
IaC tools (Ansible/T Terraform)
Metrics/logs analysis (PromQL)
End-to-end ownership
Willingness for hands-on datacenter

Job description

ElevenLabs is hiring for a remote HPC-Infrastructure Engineer specialized in GPU clusters. You’ll own GPU fleets, automate provisioning, scheduling, and monitoring, and ensure fast, reliable research compute across NVIDIA stacks and data centers.

Responsibilities include OS images, drivers, CUDA, NCCL, and high-speed networking, plus secure operations. We value hands-on engineers who thrive on end-to-end scope, enjoy building automation in Python/Bash, and can work with Slurm, Terraform, and

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Remote GPU Cluster Engineer & Automation Lead
Remote GPU Cluster Engineer & Automation Lead

AI Chopping Block • Northern (KY)

Hybrid
USD 150,000 - 230,000
HPC Infrastructure Engineer - GPU Clusters
HPC Infrastructure Engineer - GPU Clusters

AI Chopping Block • Northern (KY)

Hybrid
USD 150,000 - 230,000
HPC Infrastructure Engineer - GPU Clusters
HPC Infrastructure Engineer - GPU Clusters

ElevenLabs • Northern (KY)

Hybrid
USD 140,000 - 210,000
Remote HPC Solutions Engineer: GPU Clusters & SLURM
Remote HPC Solutions Engineer: GPU Clusters & SLURM

Hydra Host • United States

Remote
USD 120,000 - 160,000
Competitive salary
Equity
Benefits
Senior GPU Infra Engineer: AI Clusters & OpenStack Lead
Senior GPU Infra Engineer: AI Clusters & OpenStack Lead

Hamilton Barnes Associates Limited • Town of Texas (WI)

On-site
USD 120,000 - 160,000
Potential equity/bonus
Senior HPC Systems Engineer — Secure Hybrid GPU Clusters
Senior HPC Systems Engineer — Secure Hybrid GPU Clusters

Parallel Works • Chicago (IL)

Hybrid
USD 140,000 - 190,000
Medical, vision, dental coverage
401(k) with company match
Short term disability
+1
Staff Engineer, Distributed GPU Clusters
Staff Engineer, Distributed GPU Clusters

Kindredventures • San Francisco (CA)

On-site
USD 140,000 - 230,000
SRE / Platform Engineer, GPU Infrastructure
SRE / Platform Engineer, GPU Infrastructure

Bake AI • Hillsboro (OR)

On-site
USD 140,000 - 210,000
Remote HPC Solutions Engineer - GPU Clusters, ML & Equity
Remote HPC Solutions Engineer - GPU Clusters, ML & Equity

Slope • South Carolina

Remote
USD 120,000 - 160,000
Competitive salary
Equity and benefits
Flexible PTO
+1
Senior HPC Systems Engineer: GPU Clusters & AI Infra
Senior HPC Systems Engineer: GPU Clusters & AI Infra

Nebius • United States

Remote
USD 180,000 - 240,000
Competitive pay
Career growth
Flexibility and ownership
+3