Member of Technical Staff, Cluster Administration

Inferact Inc.

San Francisco, Northern (CA, KY)

Hybrid

USD 200,000 - 400,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Health benefits
Dental benefits
Vision benefits
401(k) match

Job summary

Inferact Inc. is building a world-class GPU compute platform powering vLLM. We seek a hands-on cluster administration engineer to own the high-performance infrastructure that keeps our engineers productive.

You will monitor GPU servers, manage scheduling with SLURM/Kubernetes, automate operations with Bash and Python, and respond to incidents end-to-end across multi-provider deployments. This San Francisco–based, remote-friendly role offers strong compensation and equity for engineers who thrive

Qualifications

  • Bachelor's degree or equivalent in CS, engineering, or related field.
  • Experience administering large compute clusters (HPC, GPU) and Linux.
  • Strong Linux fundamentals across networking, storage, logs, and debugging.
  • Experience with cluster scheduling using SLURM or Kubernetes.
  • Ability to own incidents end-to-end and automate workflows.

Responsibilities

  • Own and operate high-performance GPU compute infrastructure.
  • Monitor health, availability, and performance; respond to incidents.
  • Collaborate to standardize provisioning and scaling across providers.
  • Diagnose issues and optimize cluster utilization across multi-provider deployments.

Skills

Linux admin
HPC clusters
GPU servers
SLURM/Kubernetes
Automation (Bash/Python)
Incident response
Networking/storage basics

Education

Bachelor's degree

Tools

SLURM
Kubernetes
Terraform
Ansible

Job description

Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster. Founded by the creators and core maintainers of vLLM, we sit at the intersection of models and hardware—a position that took years to build.

About the Role

We're looking for a hands-on cluster administration engineer to own and operate the high-performance GPU compute infrastructure that keeps Inferact engineering productive. Inferact runs on expensive, high-performance GPU and HPC clusters across neo-cloud and dedicated compute providers. Your job is to make sure that infrastructure is healthy, available, observable, and usable around the clock.

You'll take ownership of cluster health, GPU availability, monitoring, alerting, scheduling, access, diagnostics, and incident response across the systems our engineers rely on every day. You'll work closely with engineering leadership and infrastructure owners to standardize how we provision, operate, debug, and scale compute across providers. Your work will directly impact how fast Inferact can build, test, and improve the systems powering vLLM.

Skills and Qualifications

Minimum qualifications:

  • Bachelor's degree or equivalent experience in computer science, engineering, systems administration, or similar.

  • Hands-on experience administering large compute clusters, HPC environments, university or research clusters, supercomputing systems, or production GPU clusters.

  • Strong Linux systems administration fundamentals across networking, processes, storage, package management, shell scripting, logs, access control, and system debugging.

  • Experience operating GPU servers, including driver management, GPU health monitoring, node failures, memory errors, scheduler issues, and hardware diagnostics.

  • Experience with cluster scheduling and resource allocation using SLURM, Kubernetes, or equivalent tooling.

  • Ability to own urgent infrastructure incidents end-to-end when compute issues are blocking engineering teams.

  • Ability to automate operational workflows using Bash, Python, Ansible, Terraform, Helm, or similar tooling.

Preferred qualifications:

  • Experience operating GPU compute across providers such as Lambda, CoreWeave, Crusoe, Nebius, Together, Fireworks, RunPod, or similar environments.

  • Experience improving cluster utilization, reducing idle or unavailable GPU capacity, and debugging scheduling or resource contention issues.

  • Familiarity with high-performance GPU networking such as InfiniBand, RoCE, NVLink / NVSwitch, RDMA, NCCL, or equivalent systems.

  • Experience with storage for HPC or ML workloads, including NFS, Lustre, Ceph, distributed filesystems, or other high-throughput storage systems.

  • Experience managing secure access, identity, permissions, SSH, VPNs, bastion hosts, secrets, and basic infrastructure security hygiene.

  • Background in research computing, scientific computing, ML infrastructure, SRE, platform engineering, or infrastructure operations for engineering-heavy teams.

Bonus points if you have:

  • Managed GPU or HPC infrastructure in a university lab, national lab, research institution, AI infrastructure company, hedge fund, HFT firm, or large-scale ML platform team.

  • Built monitoring, alerting, runbooks, health checks, or remediation workflows that materially reduced operational toil or incident resolution time.

  • Operated Kubernetes clusters for ML or GPU workloads at meaningful scale.

  • Standardized provisioning, diagnostics, monitoring, and operating patterns across multiple compute providers.

  • Carried real operational responsibility for infrastructure used by many engineers or researchers.

Logistics
  • Location: This role is based in San Francisco, California. Will consider remote in the US for exceptional candidates.

  • Compensation: Depending on background, skills, and experience, the expected annual salary range for this position is $200,000 - $400,000 USD + equity.

  • Visa sponsorship: We sponsor visas on a case-by-case basis.

  • Benefits: Inferact offers generous health, dental, and vision benefits as well as 401(k) company match.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Head of Engineering
Head of Engineering

Inferact • San Francisco (CA)

On-site
USD 260,000 - 380,000
Health, dental, vision benefits
401(k) company match
Member of Technical Staff, Performance and Scale
Member of Technical Staff, Performance and Scale

Inferact • San Francisco (CA)

Hybrid
USD 200,000 - 400,000
Generous health, dental, and vision benefits
401(k) company match
Equity options
Member of Technical Staff, Site Reliability Engineer
Member of Technical Staff, Site Reliability Engineer

Inferact • San Francisco (CA)

On-site
USD 200,000 - 400,000
Health benefits
401(k) match
Equity
Member of Technical Staff, Site Reliability Engineer
Member of Technical Staff, Site Reliability Engineer

Linuxconfig • San Francisco (CA), Northern (KY)

Hybrid
USD 200,000 - 400,000
Health benefits
Dental benefits
Vision benefits
+1
Senior GPU Compute Cluster Engineer
Senior GPU Compute Cluster Engineer

Inferact Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 200,000 - 400,000
Health benefits
Dental benefits
Vision benefits
+1
Member of Technical Staff, Exceptional Generalist (Remote)
Member of Technical Staff, Exceptional Generalist (Remote)

Inferact • United States

Remote
USD 180,000 - 240,000
Competitive salary and equity
Visa sponsorship
Health coverage where applicable
IT Support & Operations Engineer
IT Support & Operations Engineer

Inferact • San Francisco (CA)

On-site
USD 125,000 - 170,000
Health, dental, and vision benefits
401(k) company match
Startup equity
IT Support & Operations Engineer
IT Support & Operations Engineer

Inferact Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 125,000 - 170,000
Health benefits
Dental benefits
Vision benefits
+1
Product Marketing Manager
Product Marketing Manager

Inferact • San Francisco (CA)

On-site
USD 120,000 - 160,000
Health benefits
Dental benefits
Vision benefits
+1
Inference Infrastructure Engineer, Serving
Inference Infrastructure Engineer, Serving

Elorian • Palo Alto (CA)

On-site
USD 200,000 - 400,000
Health, dental, and vision benefits
Unlimited PTO
Parental leave
+1