Senior GPU Compute Cluster Engineer

Inferact

San Francisco (CA)

On-site

USD 200,000 - 400,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Health, dental, vision coverage
401(k) company match
Remote-friendly environment
Equity

Job summary

Inferact is seeking a hands-on cluster administration engineer to own and operate its high-performance GPU compute infrastructure. You will ensure health, availability, observability, and usability around the clock for HP compute clusters across providers.

You will collaborate with engineering leadership to standardize provisioning, operation, debugging, and scaling of compute resources. The role directly impacts speed of building and testing vLLM-powered systems.

Qualifications

  • Bachelor's degree or equivalent experience in computer science, engineering, systems administration, or similar.
  • Hands-on experience administering large compute clusters, HPC environments, university or research clusters, supercomputing systems, or production GPU clusters.
  • Strong Linux systems administration fundamentals across networking, processes, storage, package management, shell scripting, logs, access control, and system debugging.
  • Experience operating GPU servers, including driver management, GPU health monitoring, node failures, memory errors, scheduler issues, and hardware diagnostics.
  • Experience with cluster scheduling and resource allocation using SLURM, Kubernetes, or equivalent tooling.
  • Ability to own urgent infrastructure incidents end-to-end when compute issues are blocking engineering teams.
  • Ability to automate operational workflows using Bash, Python, Ansible, Terraform, Helm, or similar tooling.

Responsibilities

  • Own and operate high-performance GPU compute infrastructure used by Inferact engineers.
  • Maintain cluster health, GPU availability, monitoring, alerting, and incident response across systems.
  • Standardize provisioning and operation across providers with leadership and infra owners.
  • Work with teams to enable faster build, test, and improvement cycles for vLLM-powered systems.

Skills

Linux systems administration
GPU server handling
Cluster health monitoring
Incident response
SLURM scheduling
Kubernetes
Bash
Python
Ansible
Terraform
Helm
Remote collaboration with leadership
Hardware diagnostics

Education

Bachelor's degree in CS/Engineering or equivalent experience

Tools

SLURM
Kubernetes
Terraform
Ansible
Helm

Job description

Inferact is seeking a hands-on cluster administration engineer to own and operate its high-performance GPU compute infrastructure. You will ensure health, availability, observability, and usability around the clock for HP compute clusters across providers.

You will collaborate with engineering leadership to standardize provisioning, operation, debugging, and scaling of compute resources. The role directly impacts speed of building and testing vLLM-powered systems.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior GPU Compute Cluster Engineer
Senior GPU Compute Cluster Engineer

Inferact Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 200,000 - 400,000
Health benefits
Dental benefits
Vision benefits
+1
Member of Technical Staff, Cluster Administration
Member of Technical Staff, Cluster Administration

Inferact Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 200,000 - 400,000
Health benefits
Dental benefits
Vision benefits
+1
Senior HPC Engineer - GPU Compute & InfiniBand
Senior HPC Engineer - GPU Compute & InfiniBand

Nebius • United States

Remote
USD 150,000 - 230,000
AI Infra & Cluster Engineer — Scale GPU/CPU Orchestration
AI Infra & Cluster Engineer — Scale GPU/CPU Orchestration

Linuxcareers • San Francisco (CA)

On-site
USD 120,000 - 160,000
Senior HPC Support Engineer - GPU Cloud Infra
Senior HPC Support Engineer - GPU Cloud Infra

Neura Market • United States

On-site
USD 150,000 - 190,000
Wellness stipend
Commuter stipend
401k plan with 2% company match (USA)
+1
Senior GPU Infra Engineer for Distributed AI
Senior GPU Infra Engineer for Distributed AI

Andromeda Cluster • San Francisco (CA)

Hybrid
USD 180,000 - 240,000
Senior GPU Cloud Infrastructure Engineer
Senior GPU Cloud Infrastructure Engineer

Hyperbolic • San Francisco (CA)

On-site
USD 180,000 - 240,000
Senior GPU Cluster Engineer for AI Infrastructure
Senior GPU Cluster Engineer for AI Infrastructure

Sciforium • San Francisco (CA)

On-site
USD 150,000 - 220,000
Medical, dental, and vision insurance
401k plan
Daily lunch, snacks, and beverages
+2
HPC Cluster Engineer: Linux, InfiniBand & AI Workloads
HPC Cluster Engineer: Linux, InfiniBand & AI Workloads

Inflowfed • Springfield (VA)

On-site
USD 120,000 - 150,000
Remote GPU Cluster Engineer & Automation Lead
Remote GPU Cluster Engineer & Automation Lead

AI Chopping Block • Northern (KY)

Hybrid
USD 150,000 - 230,000