Software Engineer: ML Infra

Generalist AI

San Mateo, Somerville (CA, MA)

On-site

USD 180,000 - 240,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Generalist is building general intelligence for the physical world, operating large-scale GPU infrastructure and on-prem hardware for distributed training and robotics inference. You will own and optimize GPU fleets and data pipelines to empower researchers across workloads.

You will contribute deep expertise in ML hardware, storage, and networking, leveraging Slurm and Kubernetes to orchestrate ML tasks and robot inference fleets in distributed environments.

Qualifications

  • Own and manage GPU compute fleets for research and training.
  • Ensure GPUs are accessible and maximally utilized for researchers.
  • Build and optimize data loading and storage for distributed workloads.
  • Experience with GPU orchestration and ML infrastructure.

Responsibilities

  • Owning our GPU compute fleets
  • Ensure GPUs are easy for researchers to use and maximally utilized
  • Optimizing and improving ML data loading transport and storage in highly distributed fully utilized environments.
  • Orchestration of robot inference fleets

Skills

GPU fleet management
Kubernetes for ML
Slurm for ML
Large-scale distributed training
ML data loading optimization
NVidia GPU ecosystem

Tools

Slurm
Kubernetes
NVIDIA CUDA

Job description

About Generalist

At Generalist, we are on a mission to build general intelligence for the physical world and make it useful to everyone. We believe the industries and homes of the future will depend on humans and machines working together in new ways. Robots can help us build more and get more done. We build embodied foundation models, starting with a focus on dexterity. This requires advancing the frontiers of data, models, and hardware, to enable robots to intelligently interact with the physical world. The company embraces both large-scale AI and robotics as core to its DNA. Our team of researchers, roboticists, and company builders come from OpenAI, Boston Dynamics, Google DeepMind, and other frontier labs—with a track record of shipping AI breakthroughs. Before Generalist, we pioneered large embodied multimodal models and vision-language-action models (PaLM-E, RT-2, Gemini Robotics), launched and scaled ChatGPT and GPT-4 to hundreds of millions of users, engineered the foundations of autonomous driving, built next-generation robots (Atlas, Spot, Stretch) and pushed the limits of what they can do (from parkour to manipulation, and testing robustness). We are an equal opportunity employer, and we do not discriminate on the basis of race, religion, color, national origin, sex, sexual orientation, age, veteran status, disability, genetic information, or other applicable legally protected characteristic.

About the Role

Generalist trains very large robot foundation models. This requires utilizing very large numbers of the latest generation GPU hardware and infrastructure (currently Nvidia) to run distributed training jobs and researcher experiments. We have extreme requirements on storage and data loading infrastructure that requires maximizing cloud infrastructure and custom solutions. You will also own inference infrastructure. For our robots this is a fleet of on-prem GPUs attached to robots that have extreme real-time and latency budgets in compute constrained environments.

You’ll be responsible for:

  • Owning our GPU compute fleets
  • Ensure our GPUs are easy for researchers to use and maximally utilized
  • Optimizing and improving ML data loading transport and storage in highly distributed fully utilized environments.
  • Orchestration of robot inference fleets

You might thrive in this role if you:

  • Have managed large fleets of GPUs doing large-scale, long-term, highly distributed training runs or inference
  • Deep experience in Slurm or Kubernetes for ML workload orchestration
  • Have built high-scale ML data loaders and preparation systems
  • Deeply understand every layer of the ML hardware, storage, and networking stacks
  • Have experience in the NVidia GPU ecosystem
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Software Engineer: ML Infra
Software Engineer: ML Infra

Generalist • San Francisco (CA)

On-site
USD 120,000 - 160,000
Software Engineer: ML Robotics Systems
Software Engineer: ML Robotics Systems

Generalist AI • San Mateo (CA), Somerville (MA)

On-site
USD 170,000 - 250,000
Software Engineer: ML Optimization
Software Engineer: ML Optimization

Generalist AI • San Mateo (CA), Somerville (MA)

On-site
USD 180,000 - 240,000
Software Engineer: ML Infra
Software Engineer: ML Infra

Generalist • Somerville (MA), San Mateo (CA)

On-site
USD 120,000 - 160,000
Research Scientist: Pretraining
Research Scientist: Pretraining

Generalist AI • San Mateo (CA), Somerville (MA)

On-site
USD 180,000 - 240,000
Software Engineer: ML Optimization
Software Engineer: ML Optimization

Generalist • San Francisco (CA)

On-site
USD 120,000 - 160,000
ML Infra Engineer: GPU Fleet & Inference Orchestrator
ML Infra Engineer: GPU Fleet & Inference Orchestrator

Generalist • San Francisco (CA)

On-site
USD 120,000 - 160,000
Software Engineer: ML Optimization
Software Engineer: ML Optimization

Generalist • Somerville (MA), San Mateo (CA)

On-site
USD 100,000 - 150,000
ML Infra Engineer: GPU Fleet & Orchestration
ML Infra Engineer: GPU Fleet & Orchestration

Generalist AI • San Mateo (CA), Somerville (MA)

On-site
USD 180,000 - 240,000
Software Engineer: ML Robotics Systems
Software Engineer: ML Robotics Systems

Generalist • Somerville (MA), San Mateo (CA)

On-site
USD 100,000 - 130,000