Staff Engineer, Distributed GPU Clusters

Kindredventures

San Francisco (CA)

On-site

USD 140,000 - 230,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Kindredventures is building a Large Physics foundation Model and seeks an infrastructure engineer to design, deploy, and operate its GPU-driven compute environment. You will enable research at scale by provisioning, upgrading, and optimizing distributed clusters that power training and inference workloads.

You will extend orchestration, implement topology-aware scheduling, and deliver a self-serve platform for researchers, with strong focus on reliability, observability, and cost efficiency.

Qualifications

  • Experience operating large-scale GPU clusters and orchestration tools.
  • Strong systems background: Linux, networking, storage.
  • Knowledge of cloud platforms (GCP, AWS, or Azure).
  • Understanding of monitoring, logging, observability, and version control for ML systems.
  • Familiarity with CUDA/NCCL for distributed workloads.
  • Owns deliverables end-to-end, from requirements through autonomous execution.

Responsibilities

  • Design, deploy, and operate large distributed GPU clusters end to end: provisioning, imaging, upgrades, and capacity planning.
  • Extend scheduling and orchestration systems for topology-aware placement, preemption, quotas, and multi-tenancy across training and inference workloads.
  • Build software that abstracts cluster management and presents a unified, self-serve interface to researchers and engineers.
  • Own cluster storage and artifact paths for checkpoints and logs, with clear retention and lineage.
  • Monitor and continuously improve reliability and error recovery; build observability to catch failures before researchers do.
  • Partner with researchers to unblock large-scale runs and advise on performance and placement trade-offs.

Skills

GPU clusters
Container orchestration
Systems background
Cloud platforms
Monitoring observability
CUDA/NCCL
End-to-end ownership

Tools

Kubernetes
Slurm
Docker

Job description

Kindredventures is building a Large Physics foundation Model and seeks an infrastructure engineer to design, deploy, and operate its GPU-driven compute environment. You will enable research at scale by provisioning, upgrading, and optimizing distributed clusters that power training and inference workloads.

You will extend orchestration, implement topology-aware scheduling, and deliver a self-serve platform for researchers, with strong focus on reliability, observability, and cost efficiency.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff Engineer, Distributed GPU Clusters & Infra
Staff Engineer, Distributed GPU Clusters & Infra

Causal Labs • San Francisco (CA)

On-site
USD 180,000 - 240,000
Staff GPU Compute Infrastructure Engineer
Staff GPU Compute Infrastructure Engineer

Causal • San Francisco (CA)

On-site
USD 180,000 - 240,000
GPU Systems Engineer
GPU Systems Engineer

Career Techniques • New York (NY)

Hybrid
USD 200,000 - 300,000
Senior GPU Infra Engineer for Distributed AI
Senior GPU Infra Engineer for Distributed AI

Andromeda Cluster • San Francisco (CA)

Hybrid
USD 180,000 - 240,000
Member of Technical Staff — Compute Cluster
Member of Technical Staff — Compute Cluster

Linuxcareers • San Francisco (CA)

On-site
USD 120,000 - 180,000
Distributed Training Infra Engineer for Large Models
Distributed Training Infra Engineer for Large Models

Kindredventures • San Francisco (CA)

On-site
USD 180,000 - 240,000
Member of Technical Staff — Compute Cluster
Member of Technical Staff — Compute Cluster

Causal Labs • San Francisco (CA)

On-site
USD 180,000 - 240,000
GPU Cluster Infrastructure Engineer
GPU Cluster Infrastructure Engineer

PVH (Tommy Hilfiger/Calvin Klein) • United States

On-site
USD 140,000 - 210,000
Equity in unicorn-stage company
100% premiums covered for medical, den
401(k) matching up to 4%
+2
Staff Inference Systems Engineer — High-Throughput AI
Staff Inference Systems Engineer — High-Throughput AI

Kindredventures • San Francisco (CA)

On-site
USD 180,000 - 240,000
Member of Technical Staff — Compute Cluster
Member of Technical Staff — Compute Cluster

Kindredventures • San Francisco (CA)

On-site
USD 140,000 - 230,000