Software Engineer: ML Infra

Generalist

Somerville, San Mateo (MA, CA)

On-site

USD 120,000 - 160,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Generalist is seeking a talented individual in Somerville, Massachusetts to manage GPU fleets for large-scale robot training. The role involves optimizing data loading systems and orchestrating inference fleets, ensuring high performance under compute constraints.

The ideal candidate has deep experience with Slurm or Kubernetes for machine learning orchestration and a solid understanding of the Nvidia GPU ecosystem. Come join an equal opportunity employer that values diversity!

Qualifications

  • Experience managing large fleets of GPUs for distributed training.
  • Deep experience with Slurm or Kubernetes for ML workloads.
  • Knowledge in building high-scale ML data loaders.

Responsibilities

  • Own the GPU compute fleets and ensure maximal utilization.
  • Optimize ML data loading transport and storage.
  • Orchestrate robot inference fleets.

Skills

GPU resource management
ML workload orchestration (Slurm or Kubernetes)
Data loading optimization
Understanding of ML hardware
Experience in NVidia GPU ecosystem

Job description

About the Role

Generalist trains very large robot foundation models. This requires utilizing very large numbers of the latest generation GPU hardware and infrastructure (currently Nvidia) to run distributed training jobs and researcher experiments. We have extreme requirements on storage and data loading infrastructure that requires maximizing cloud infrastructure and custom solutions.

You will also own inference infrastructure. For our robots this is a fleet of on-prem GPUs attached to robots that have extreme real-time and latency budgets in compute constrained environments.

You’ll be responsible for:
  • Owning our GPU compute fleets
  • Ensure our GPUs are easy for researchers to use and maximally utilized
  • Optimizing and improving ML data loading transport and storage in highly distributed fully utilized environments.
  • Orchestration of robot inference fleets
You might thrive in this role if you:
  • Have managed large fleets of GPUs doing large-scale, long-term, highly distributed training runs or inference
  • Deep experience in Slurm or Kubernetes for ML workload orchestration
  • Have build high-scale ML data loaders and preparation systems
  • Deeply understand every layer of the ML hardware, storage, and networking stacks
  • Have experience in the NVidia GPU ecosystem

We are an equal opportunity employer, and we do not discriminate on the basis of race, religion, color, national origin, sex, sexual orientation, age, veteran status, disability, genetic information, or other applicable legally protected characteristic.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Software Engineer: ML Infra
Software Engineer: ML Infra

Generalist • San Francisco (CA)

On-site
USD 120,000 - 160,000
ML Infra Engineer: GPU Fleet & Inference Orchestrator
ML Infra Engineer: GPU Fleet & Inference Orchestrator

Generalist • San Francisco (CA)

On-site
USD 120,000 - 160,000
Member of Technical Staff (Software Engineer, Inference & Training Platform)
Member of Technical Staff (Software Engineer, Inference & Training Platform)

United States Digital Space LLC • San Francisco (CA)

Hybrid
USD 180,000 - 240,000
Software Engineer, ML Infrastructure
Software Engineer, ML Infrastructure

Cursor • New York (NY), San Francisco (CA)

On-site
USD 120,000 - 150,000
Member of Technical Staff - ML Infra
Member of Technical Staff - ML Infra

Kindredventures • San Francisco (CA)

On-site
USD 160,000 - 220,000
Senior ML Infra Engineer - Large-Scale Training & Pipelines
Senior ML Infra Engineer - Large-Scale Training & Pipelines

Kindredventures • San Francisco (CA)

On-site
USD 160,000 - 220,000
Principal ML Infrastructure Engineer (Relocation Available)
Principal ML Infrastructure Engineer (Relocation Available)

Franklin Fitch • Dallas (TX)

On-site
USD 100,000 - 140,000
Member of Technical Staff – Software Engineer, GPU Cluster Infrastructure
Member of Technical Staff – Software Engineer, GPU Cluster Infrastructure

Perplexity • San Francisco (CA)

On-site
USD 180,000 - 240,000
AI/ML Infra Engineer - Hosting
AI/ML Infra Engineer - Hosting

Hamilton Barnes Associates Limited • San Francisco (CA)

On-site
USD 225,000 - 275,000
Stock options
Senior Software Engineer, Infrastructure Software for AI (Centralized AI Data Centers & Distrib[...]
Senior Software Engineer, Infrastructure Software for AI (Centralized AI Data Centers & Distrib[...]

Intelliswift - An LTTS Company • Sunnyvale (CA)

On-site
USD 120,000 - 150,000
Competitive salary
Health insurance
Flexible work hours