Senior ML Infra Engineer - GPU Compute Leader

Flexion Robotics

San Francisco (CA)

On-site

USD 180,000 - 240,000

Full time

3 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Competitive Compensation
Leading robotics team
Energetic culture

Job summary

Flexion Robotics seeks a senior ML engineer to own GPU compute platforms and optimize training workflows on multi-node clusters. You will collaborate with AI and infrastructure teams to improve speed, utilization, and scalability across on-prem and cloud environments.

The role requires deep knowledge of distributed training, PyTorch, and Python, with hands-on experience in large-model training and performance profiling. This is an on-site opportunity in San Francisco with leadership exposure.

Qualifications

  • Degree in CS/EE/Software Eng or equivalent practical experience.
  • Hands-on with training/inference of large models on distributed multi-node GPU hardware.
  • Proficiency in Python and working knowledge of PyTorch.
  • Deep understanding of distributed training concepts (DDP, FSDP, NCCL).
  • Experience with at least one cloud platform (AWS, GCP, Azure) or large-scale on-premises GPU infra.
  • Experience with job scheduling/orchestration tools: Slurm and/or Kubernetes.

Responsibilities

  • Architect and run cloud-based GPU clusters; select frameworks and tooling.
  • Help AI engineers optimize training workloads and hardware utilization.
  • Contribute to GPU compute strategies with AI engineering teams.
  • Explore multi-cloud strategies to optimize capacity and cost.
  • Raise engineering quality with testing, docs, and reliability.

Skills

Python
PyTorch
Distributed training
Cloud platforms
Slurm/Kubernetes
Large models

Education

Degree in Computer Science / Electrical Engineering / Software Engineering

Tools

Kubernetes
Slurm
Terraform
Ansible

Job description

Flexion Robotics seeks a senior ML engineer to own GPU compute platforms and optimize training workflows on multi-node clusters. You will collaborate with AI and infrastructure teams to improve speed, utilization, and scalability across on-prem and cloud environments.

The role requires deep knowledge of distributed training, PyTorch, and Python, with hands-on experience in large-model training and performance profiling. This is an on-site opportunity in San Francisco with leadership exposure.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior ML Infra Engineer — GPU Clusters for Humanoid AI
Senior ML Infra Engineer — GPU Clusters for Humanoid AI

Flexion • San Francisco (CA)

On-site
USD 180,000 - 240,000
Competitive Compensation
Relocation sponsorship
401(k) with company contributions
+2
ML Engineer - Infrastructure
ML Engineer - Infrastructure

Flexion Robotics • San Francisco (CA)

On-site
USD 180,000 - 240,000
Competitive Compensation
Leading robotics team
Energetic culture
ML Engineer - Infrastructure
ML Engineer - Infrastructure

Flexion • San Francisco (CA)

On-site
USD 180,000 - 240,000
Competitive Compensation
Relocation sponsorship
401(k) with company contributions
+2
Senior ML Infra Engineer - Scale GPU Clusters, Remote
Senior ML Infra Engineer - Scale GPU Clusters, Remote

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 320,000 - 500,000
Equity
Medical/Dental/Vision coverage
Unlimited PTO
+1
Senior ML Infrastructure Engineer - GPU Training & MLOps
Senior ML Infrastructure Engineer - GPU Training & MLOps

Atoms • San Francisco (CA)

On-site
USD 224,000 - 280,000
Medical, Dental, Vision, Disability, and Life Insurance
Flexible Spending Account / Health Savings Account Options
401(k)
+2
ML Infra Engineer: GPU Fleet & Inference Orchestrator
ML Infra Engineer: GPU Fleet & Inference Orchestrator

Generalist • San Francisco (CA)

On-site
USD 120,000 - 160,000
ML Infrastructure Engineer — On-Site SF, GPU Pipelines
ML Infrastructure Engineer — On-Site SF, GPU Pipelines

Objective Partners • San Francisco (CA)

On-site
USD 180,000 - 250,000
Full medical, dental, vision coverage
Flexible PTO
Daily catered lunches
+1
ML Infra Engineer — GPU Clusters & Distributed Systems
ML Infra Engineer — GPU Clusters & Distributed Systems

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 170,000 - 250,000
Industry-leading compensation and/or:?
Unlimited PTO
Top-tier medical, dental, and vision
+1
Software Engineer: ML Infra
Software Engineer: ML Infra

Generalist • Somerville (MA), San Mateo (CA)

On-site
USD 120,000 - 160,000
Senior ML Training Systems Engineer - Distributed GPU Infra
Senior ML Training Systems Engineer - Distributed GPU Infra

Baseten • San Francisco (CA)

On-site
USD 150,000 - 200,000
Competitive compensation, including equity
100% coverage of medical, dental, and vision insurance
Generous PTO policy
+2