Senior ML Infra Engineer — GPU Clusters for Humanoid AI

Flexion

San Francisco (CA)

On-site

USD 180,000 - 240,000

Full time

14 hours ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Competitive Compensation
Relocation sponsorship
401(k) with company contributions
Health, dental & vision coverage
Open PTO

Job summary

Flexion seeks a senior ML engineer to own its GPU compute platforms on-site in San Francisco. You will design, bring up, operate and optimize distributed multi-node GPU clusters and collaborate with AI engineers to accelerate training and improve hardware utilization.

You will influence long-term compute strategy with the infrastructure and AI teams, explore multi-cloud options, and raise engineering standards with testing, docs, and reliability practices.

Qualifications

  • Hands-on experience with the training or inference of large models (billions of parameters) on distributed multi-node GPU hardware.
  • Proficiency in Python and working knowledge of PyTorch.
  • Deep understanding of distributed training concepts (DDP, FSDP, NCCL).
  • Experience with at least one cloud platform (AWS, GCP, Azure or neoclouds) or large-scale on-premises GPU infrastructure.
  • Experience with job scheduling and orchestration tools: Slurm and/or Kubernetes/KubeRay.

Responsibilities

  • Architect, run and continuously improve existing and future cloud-based GPU clusters. Select the best frameworks and tooling to run our clusters efficiently. Work on cluster provisioning, job schedulers and monitoring systems
  • Help AI engineers optimize their training workloads and maximize hardware utilization using profilers, contributing to our core ML libraries
  • Contribute to short- and long-term GPU compute strategies in collaboration with our AI engineering teams and help execute on them.
  • Optimize capacity and cost by exploring multi-cloud strategies and evaluating trade-offs
  • Raise the bar on engineering practices, including testing, code quality, documentation, and system reliability

Skills

Python
PyTorch
Distributed training
Slurm
Kubernetes
AWS/GCP/Azure

Education

Degree in CS or related

Tools

KubeRay
Profilers

Job description

Flexion seeks a senior ML engineer to own its GPU compute platforms on-site in San Francisco. You will design, bring up, operate and optimize distributed multi-node GPU clusters and collaborate with AI engineers to accelerate training and improve hardware utilization.

You will influence long-term compute strategy with the infrastructure and AI teams, explore multi-cloud options, and raise engineering standards with testing, docs, and reliability practices.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior ML Infra Engineer - GPU Compute Leader
Senior ML Infra Engineer - GPU Compute Leader

Flexion Robotics • San Francisco (CA)

On-site
USD 180,000 - 240,000
Competitive Compensation
Leading robotics team
Energetic culture
ML Engineer - Infrastructure
ML Engineer - Infrastructure

Flexion Robotics • San Francisco (CA)

On-site
USD 180,000 - 240,000
Competitive Compensation
Leading robotics team
Energetic culture
ML Engineer - Infrastructure
ML Engineer - Infrastructure

Flexion • San Francisco (CA)

On-site
USD 180,000 - 240,000
Competitive Compensation
Relocation sponsorship
401(k) with company contributions
+2
Senior ML Infra Engineer - Scale GPU Clusters, Remote
Senior ML Infra Engineer - Scale GPU Clusters, Remote

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 320,000 - 500,000
Equity
Medical/Dental/Vision coverage
Unlimited PTO
+1
ML Infrastructure Engineer — On-Site SF, GPU Pipelines
ML Infrastructure Engineer — On-Site SF, GPU Pipelines

Objective Partners • San Francisco (CA)

On-site
USD 180,000 - 250,000
Full medical, dental, vision coverage
Flexible PTO
Daily catered lunches
+1
Senior GPU Infrastructure Engineer — HPC & Clusters
Senior GPU Infrastructure Engineer — HPC & Clusters

Prime Intellect AI • San Francisco (CA)

On-site
USD 150,000 - 300,000
ML Infra Engineer — GPU Clusters & Distributed Systems
ML Infra Engineer — GPU Clusters & Distributed Systems

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 170,000 - 250,000
Industry-leading compensation and/or:?
Unlimited PTO
Top-tier medical, dental, and vision
+1
Senior GPU Cluster Architect for AI Infra at Scale
Senior GPU Cluster Architect for AI Infra at Scale

Partner Company • United States

Remote
USD 184,000 - 318,000
Medical insurance
Dental insurance
Vision insurance
+1
Senior ML Training Systems Engineer - Distributed CUDA
Senior ML Training Systems Engineer - Distributed CUDA

Genesis AI • San Francisco (CA)

On-site
USD 180,000 - 260,000
Senior AI Infrastructure Lead: GPU Clusters & LLMs
Senior AI Infrastructure Lead: GPU Clusters & LLMs

Cadence Design Systems • San Jose (CA)

On-site
USD 137,000 - 254,000