Member of Technical Staff — Compute Cluster

Linuxcareers

San Francisco (CA)

On-site

USD 120,000 - 180,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Linuxcareers in San Francisco is building AI research infrastructure. You will design, deploy, and operate large-scale GPU clusters powering training, evaluation, and serving for the research team.

The role emphasizes extending orchestration with Kubernetes/Slurm, building unified interfaces, and ensuring reliability with observability. You'll work with researchers to optimize performance and placement.

Qualifications

  • Experience operating large-scale GPU clusters.
  • Proficient with Kubernetes, Slurm and Docker.
  • Strong Linux, networking, storage, and IaC skills.
  • Familiar with cloud ML services (GCP/AWS/Azure).
  • CUDA/NCCL experience and profiling for distributed workloads.
  • Ability to deliver end-to-end from requirements through autonomous execution.

Responsibilities

  • Design, deploy, and operate large distributed GPU clusters end-to-end, including provisioning, imaging, upgrades, and capacity planning
  • Extend scheduling and orchestration systems such as Kubernetes and Slurm for topology-aware placement, preemption, quotas, and multi-tenancy across training and inference workloads
  • Build software that abstracts cluster management and presents a unified, self-serve interface to researchers and engineers
  • Own cluster storage and artifact paths for checkpoints and logs, with clear retention and lineage
  • Monitor and continuously improve reliability and error recovery; build observability to catch failures before researchers do
  • Partner with researchers to unblock large-scale runs and advise on performance and placement trade-offs

Skills

GPU clusters
Kubernetes
Slurm
Docker
Linux systems
Networking
Storage
Infrastructure as code
Cloud platforms
CUDA/NCCL
Observability
End-to-end ownership

Tools

Kubernetes
Slurm
Docker

Job description

This AI research company is building large physics foundation models for causal intelligence and weather prediction. You will design, build, and operate large-scale GPU clusters that power training, evaluation, and serving infrastructure for the research team.

What You\'ll Do
  • Design, deploy, and operate large distributed GPU clusters end-to-end, including provisioning, imaging, upgrades, and capacity planning
  • Extend scheduling and orchestration systems such as Kubernetes and Slurm for topology-aware placement, preemption, quotas, and multi-tenancy across training and inference workloads
  • Build software that abstracts cluster management and presents a unified, self-serve interface to researchers and engineers
  • Own cluster storage and artifact paths for checkpoints and logs, with clear retention and lineage
  • Monitor and continuously improve reliability and error recovery; build observability to catch failures before researchers do
  • Partner with researchers to unblock large-scale runs and advise on performance and placement trade-offs
What You Need
  • Experience operating large-scale GPU clusters and container orchestration frameworks such as Kubernetes, Slurm, and Docker
  • Strong systems background in Linux, networking, storage, and infrastructure-as-code
  • Knowledge of cloud platforms including GCP, AWS, or Azure and their ML/AI service offerings
  • Understanding of monitoring, logging, observability, and version control best practices for ML systems
  • Familiarity with CUDA and NCCL, and performance profiling for distributed workloads
  • Ability to own deliverables end-to-end from requirements through autonomous execution
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Member of Technical Staff — Compute Cluster
Member of Technical Staff — Compute Cluster

Causal • San Francisco (CA)

On-site
USD 180,000 - 240,000
Member of Technical Staff — Compute Cluster
Member of Technical Staff — Compute Cluster

Causal Labs • San Francisco (CA)

On-site
USD 180,000 - 240,000
Member of Technical Staff — Compute Cluster
Member of Technical Staff — Compute Cluster

Kindredventures • San Francisco (CA)

On-site
USD 140,000 - 230,000
Cluster Engineer
Cluster Engineer

STN Inc • San Francisco (CA)

On-site
USD 180,000 - 240,000
Cluster Design
Cluster Design

Blue Signal Search • San Francisco (CA)

On-site
USD 150,000 - 230,000
GPU Systems Engineer
GPU Systems Engineer

Career Techniques • New York (NY)

Hybrid
USD 200,000 - 300,000
Infrastructure / Cluster Engineer
Infrastructure / Cluster Engineer

Linuxcareers • San Francisco (CA)

On-site
USD 120,000 - 160,000
Staff Engineer, Distributed GPU Clusters & Infra
Staff Engineer, Distributed GPU Clusters & Infra

Causal Labs • San Francisco (CA)

On-site
USD 180,000 - 240,000
Head of AI Data Center Infrastructure Platforms and Software
Head of AI Data Center Infrastructure Platforms and Software

Summit Group Solutions, LLC • United States

On-site
USD 150,000 - 350,000
Software Engineer, Compute Foundations
Software Engineer, Compute Foundations

Linuxcareers • San Francisco (CA), Northern (KY)

Hybrid
USD 210,000 - 270,000