ML Infra Engineer: GPU Fleet & Orchestration

Generalist AI

San Mateo, Somerville (CA, MA)

On-site

USD 180,000 - 240,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Generalist is building general intelligence for the physical world, operating large-scale GPU infrastructure and on-prem hardware for distributed training and robotics inference. You will own and optimize GPU fleets and data pipelines to empower researchers across workloads.

You will contribute deep expertise in ML hardware, storage, and networking, leveraging Slurm and Kubernetes to orchestrate ML tasks and robot inference fleets in distributed environments.

Qualifications

  • Own and manage GPU compute fleets for research and training.
  • Ensure GPUs are accessible and maximally utilized for researchers.
  • Build and optimize data loading and storage for distributed workloads.
  • Experience with GPU orchestration and ML infrastructure.

Responsibilities

  • Owning our GPU compute fleets
  • Ensure GPUs are easy for researchers to use and maximally utilized
  • Optimizing and improving ML data loading transport and storage in highly distributed fully utilized environments.
  • Orchestration of robot inference fleets

Skills

GPU fleet management
Kubernetes for ML
Slurm for ML
Large-scale distributed training
ML data loading optimization
NVidia GPU ecosystem

Tools

Slurm
Kubernetes
NVIDIA CUDA

Job description

Generalist is building general intelligence for the physical world, operating large-scale GPU infrastructure and on-prem hardware for distributed training and robotics inference. You will own and optimize GPU fleets and data pipelines to empower researchers across workloads.

You will contribute deep expertise in ML hardware, storage, and networking, leveraging Slurm and Kubernetes to orchestrate ML tasks and robot inference fleets in distributed environments.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

ML Infra Engineer: GPU Fleet & Inference Orchestrator
ML Infra Engineer: GPU Fleet & Inference Orchestrator

Generalist • San Francisco (CA)

On-site
USD 120,000 - 160,000
ML Infra & GPU Fleet Engineer
ML Infra & GPU Fleet Engineer

Generalist • Somerville (MA), San Mateo (CA)

On-site
USD 120,000 - 160,000
Software Engineer: ML Infra
Software Engineer: ML Infra

Generalist AI • San Mateo (CA), Somerville (MA)

On-site
USD 180,000 - 240,000
ML Infra Engineer: GPU Orchestration & Observability
ML Infra Engineer: GPU Orchestration & Observability

Autolab • San Francisco (CA)

On-site
USD 150,000 - 210,000
Software Engineer: ML Infra
Software Engineer: ML Infra

Generalist • Somerville (MA), San Mateo (CA)

On-site
USD 120,000 - 160,000
Software Engineer: ML Infra
Software Engineer: ML Infra

Generalist • San Francisco (CA)

On-site
USD 120,000 - 160,000
ML Infra Engineer — GPU Clusters & Distributed Systems
ML Infra Engineer — GPU Clusters & Distributed Systems

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 170,000 - 250,000
Industry-leading compensation and/or:?
Unlimited PTO
Top-tier medical, dental, and vision
+1
ML Inference Infrastructure Architect
ML Inference Infrastructure Architect

Morph • California (MO)

On-site
USD 90,000 - 120,000
Senior Remote ML Infrastructure Engineer: GPU & Scale
Senior Remote ML Infrastructure Engineer: GPU & Scale

Bright Vision Technologies • Bellevue (WA)

On-site
USD 100,000 - 150,000
Research Engineer: High-Performance ML Infrastructure
Research Engineer: High-Performance ML Infrastructure

Fleet AI, Inc. • San Francisco (CA)

On-site
USD 180,000 - 240,000