Senior GPU Compute Cluster Engineer

Inferact Inc.

San Francisco, Northern (CA, KY)

Hybrid

USD 200,000 - 400,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Health benefits
Dental benefits
Vision benefits
401(k) match

Job summary

Inferact Inc. is building a world-class GPU compute platform powering vLLM. We seek a hands-on cluster administration engineer to own the high-performance infrastructure that keeps our engineers productive.

You will monitor GPU servers, manage scheduling with SLURM/Kubernetes, automate operations with Bash and Python, and respond to incidents end-to-end across multi-provider deployments. This San Francisco–based, remote-friendly role offers strong compensation and equity for engineers who thrive

Qualifications

  • Bachelor's degree or equivalent in CS, engineering, or related field.
  • Experience administering large compute clusters (HPC, GPU) and Linux.
  • Strong Linux fundamentals across networking, storage, logs, and debugging.
  • Experience with cluster scheduling using SLURM or Kubernetes.
  • Ability to own incidents end-to-end and automate workflows.

Responsibilities

  • Own and operate high-performance GPU compute infrastructure.
  • Monitor health, availability, and performance; respond to incidents.
  • Collaborate to standardize provisioning and scaling across providers.
  • Diagnose issues and optimize cluster utilization across multi-provider deployments.

Skills

Linux admin
HPC clusters
GPU servers
SLURM/Kubernetes
Automation (Bash/Python)
Incident response
Networking/storage basics

Education

Bachelor's degree

Tools

SLURM
Kubernetes
Terraform
Ansible

Job description

Inferact Inc. is building a world-class GPU compute platform powering vLLM. We seek a hands-on cluster administration engineer to own the high-performance infrastructure that keeps our engineers productive.

You will monitor GPU servers, manage scheduling with SLURM/Kubernetes, automate operations with Bash and Python, and respond to incidents end-to-end across multi-provider deployments. This San Francisco–based, remote-friendly role offers strong compensation and equity for engineers who thrive

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior GPU Compute Cluster Engineer
Senior GPU Compute Cluster Engineer

Inferact • San Francisco (CA)

On-site
USD 200,000 - 400,000
Health, dental, vision coverage
401(k) company match
Remote-friendly environment
+1
Member of Technical Staff, Cluster Administration
Member of Technical Staff, Cluster Administration

Inferact Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 200,000 - 400,000
Health benefits
Dental benefits
Vision benefits
+1
Senior HPC & GPU Cluster Architect — Scale & Automate
Senior HPC & GPU Cluster Architect — Scale & Automate

San Francisco Compute Company • San Francisco (CA)

Hybrid
USD 120,000 - 160,000
Generous equity grant
Competitive salary
Visa sponsorship
+6
Senior GPU Cluster Engineer for AI Infrastructure
Senior GPU Cluster Engineer for AI Infrastructure

Sciforium • San Francisco (CA)

On-site
USD 150,000 - 220,000
Medical, dental, and vision insurance
401k plan
Daily lunch, snacks, and beverages
+2
Head of AI Data Center Infrastructure Platforms and Software
Head of AI Data Center Infrastructure Platforms and Software

Summit Group Solutions, LLC • United States

On-site
USD 150,000 - 350,000
Senior GPU Infra Engineer for Distributed AI
Senior GPU Infra Engineer for Distributed AI

Andromeda Cluster • San Francisco (CA)

Hybrid
USD 180,000 - 240,000
GPU Infra Solutions Architect for Large-Scale AI Clusters
GPU Infra Solutions Architect for Large-Scale AI Clusters

Prime Intellect • San Francisco (CA)

On-site
USD 150,000 - 300,000
Senior GPU Infra Engineer: AI Clusters & OpenStack Lead
Senior GPU Infra Engineer: AI Clusters & OpenStack Lead

Hamilton Barnes Associates Limited • Town of Texas (WI)

On-site
USD 120,000 - 160,000
Potential equity/bonus
AI Infra & Cluster Engineer — Scale GPU/CPU Orchestration
AI Infra & Cluster Engineer — Scale GPU/CPU Orchestration

Linuxcareers • San Francisco (CA)

On-site
USD 120,000 - 160,000
Senior GPU Infra SRE — Remote, Low-Level Linux & Scale
Senior GPU Infra SRE — Remote, Low-Level Linux & Scale

Luma • Redwood City (CA)

Remote
USD 180,000 - 240,000