Senior ML Training Systems Engineer - Distributed GPU Infra

Baseten

San Francisco (CA)

On-site

USD 150,000 - 200,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Competitive compensation, including equity
100% coverage of medical, dental, and vision insurance
Generous PTO policy
Paid parental leave
Company-facilitated 401(k)

Job summary

A leading AI technology company in San Francisco is looking for a Senior Software Engineer to build scalable infrastructure for large‑scale training and fine-tuning of foundation models. You will design distributed training systems and optimize GPU utilization while collaborating with cross-functional teams to adapt models efficiently. Ideal candidates have over 5 years of experience in ML infrastructure and a strong background in distributed training frameworks. Competitive compensation and benefits package is offered.

Qualifications

  • 5+ years of experience in ML infrastructure or distributed systems.
  • Strong expertise in distributed training frameworks like FSDP and DDP.
  • Hands-on experience with LLMs or foundation models.

Responsibilities

  • Design and maintain the distributed training infrastructure.
  • Implement scalable training pipelines for GPU clusters.
  • Optimize training performance using techniques like mixed precision.

Skills

Distributed training frameworks
Kubernetes
ML infrastructure
Fine-tuning techniques
GPU optimization

Education

Bachelor’s degree in Computer Science

Tools

Ray
Slurm

Job description

A leading AI technology company in San Francisco is looking for a Senior Software Engineer to build scalable infrastructure for large‑scale training and fine-tuning of foundation models. You will design distributed training systems and optimize GPU utilization while collaborating with cross-functional teams to adapt models efficiently. Ideal candidates have over 5 years of experience in ML infrastructure and a strong background in distributed training frameworks. Competitive compensation and benefits package is offered.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior ML Infrastructure Engineer - GPU & Scale
Senior ML Infrastructure Engineer - GPU & Scale

TensorWave • Las Vegas (NV)

On-site
USD 120,000 - 150,000
Competitive Salary
Stock Options
100% paid Medical, Dental, and Vision insurance
+8
Senior GPU ML Infra Engineer — Mid-Training & Inference
Senior GPU ML Infra Engineer — Mid-Training & Inference

Reflection AI • San Francisco (CA)

On-site
USD 120,000 - 160,000
Top-tier compensation
Comprehensive medical, dental, vision, life, and disability insurance
Fully paid parental leave for all new parents
+2
Staff, Pre-Training Infra — Distributed ML Training
Staff, Pre-Training Infra — Distributed ML Training

B Capital • San Francisco (CA)

On-site
USD 120,000 - 160,000
Top-tier compensation
Comprehensive medical, dental, vision, life, and disability insurance
Fully paid parental leave
+2
Staff ML Systems Engineer — Distributed Training at Scale
Staff ML Systems Engineer — Distributed Training at Scale

RadixArk • Palo Alto (CA)

On-site
USD 120,000 - 160,000
Comprehensive benefits
Flexible work arrangements
Senior ML Training Systems Engineer - Distributed CUDA
Senior ML Training Systems Engineer - Distributed CUDA

Genesis AI • San Francisco (CA)

On-site
USD 180,000 - 260,000
Senior ML Systems Engineer: Scalable Training Frameworks
Senior ML Systems Engineer: Scalable Training Frameworks

Cohere • San Francisco (CA)

Hybrid
USD 150,000 - 180,000
Inclusive culture
Weekly lunch stipend
Health and dental benefits
+4
Senior ML Infra Architect — Large-Scale GPU Training
Senior ML Infra Architect — Large-Scale GPU Training

Hark, Inc. • San Jose (CA)

On-site
USD 180,000 - 450,000
Senior ML Performance Engineer - Distributed Training
Senior ML Performance Engineer - Distributed Training

Odyssey • Santa Clara (CA)

On-site
USD 120,000 - 160,000
Senior ML Infra Engineer - Large-Scale Training & Pipelines
Senior ML Infra Engineer - Large-Scale Training & Pipelines

Kindredventures • San Francisco (CA)

On-site
USD 160,000 - 220,000
Senior ML Infrastructure Engineer - GPU Training & MLOps
Senior ML Infrastructure Engineer - GPU Training & MLOps

Atoms • San Francisco (CA)

On-site
USD 224,000 - 280,000
Medical, Dental, Vision, Disability, and Life Insurance
Flexible Spending Account / Health Savings Account Options
401(k)
+2