Senior ML Training Systems Engineer - Distributed GPU Infra

Baseten

San Francisco (CA)

On-site

USD 150,000 - 200,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Competitive compensation, including equity
100% coverage of medical, dental, and vision insurance
Generous PTO policy
Paid parental leave
Company-facilitated 401(k)

Job summary

A leading AI technology company in San Francisco is looking for a Senior Software Engineer to build scalable infrastructure for large‑scale training and fine-tuning of foundation models. You will design distributed training systems and optimize GPU utilization while collaborating with cross-functional teams to adapt models efficiently. Ideal candidates have over 5 years of experience in ML infrastructure and a strong background in distributed training frameworks. Competitive compensation and benefits package is offered.

Qualifications

  • 5+ years of experience in ML infrastructure or distributed systems.
  • Strong expertise in distributed training frameworks like FSDP and DDP.
  • Hands-on experience with LLMs or foundation models.

Responsibilities

  • Design and maintain the distributed training infrastructure.
  • Implement scalable training pipelines for GPU clusters.
  • Optimize training performance using techniques like mixed precision.

Skills

Distributed training frameworks
Kubernetes
ML infrastructure
Fine-tuning techniques
GPU optimization

Education

Bachelor’s degree in Computer Science

Tools

Ray
Slurm

Job description

A leading AI technology company in San Francisco is looking for a Senior Software Engineer to build scalable infrastructure for large‑scale training and fine-tuning of foundation models. You will design distributed training systems and optimize GPU utilization while collaborating with cross-functional teams to adapt models efficiently. Ideal candidates have over 5 years of experience in ML infrastructure and a strong background in distributed training frameworks. Competitive compensation and benefits package is offered.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff Engineer, Distributed Pre-Training Infra
Staff Engineer, Distributed Pre-Training Infra

Reflection • San Francisco (CA)

On-site
Top-tier compensation
Comprehensive health and wellness benefits
Paid parental leave
+2
Senior ML Infrastructure Engineer - GPU & Scale
Senior ML Infrastructure Engineer - GPU & Scale

TensorWave • Las Vegas (NV)

On-site
USD 120,000 - 150,000
Competitive Salary
Stock Options
100% paid Medical, Dental, and Vision insurance
+8
Staff Engineer, Mid-Training Infra for Large-Scale AI
Staff Engineer, Mid-Training Infra for Large-Scale AI

Reflection • San Francisco (CA)

On-site
Top-tier compensation
Comprehensive health, dental, and vision insurance
Fully paid parental leave
+2
Senior GPU ML Infra Engineer — Mid-Training & Inference
Senior GPU ML Infra Engineer — Mid-Training & Inference

Reflection AI • San Francisco (CA)

On-site
USD 120,000 - 160,000
Staff, Pre-Training Infra — Distributed ML Training
Staff, Pre-Training Infra — Distributed ML Training

B Capital • San Francisco (CA)

On-site
USD 120,000 - 160,000
Top-tier compensation
Comprehensive medical, dental, vision, life, and disability insurance
Fully paid parental leave
+2
Staff ML Systems Engineer — Distributed Training at Scale
Staff ML Systems Engineer — Distributed Training at Scale

RadixArk • Palo Alto (CA)

On-site
USD 120,000 - 160,000
Comprehensive benefits
Flexible work arrangements
Senior ML Systems Engineer: Scalable Training Frameworks
Senior ML Systems Engineer: Scalable Training Frameworks

Cohere • San Francisco (CA)

Hybrid
USD 150,000 - 180,000
Senior ML Training Systems Engineer - Distributed CUDA
Senior ML Training Systems Engineer - Distributed CUDA

Genesis AI • San Francisco (CA)

On-site
USD 180,000 - 260,000
Senior Distributed ML Systems Engineer (Remote • Equity)
Senior Distributed ML Systems Engineer (Remote • Equity)

Pluralis Research • San Francisco (CA)

Remote
Equity-heavy compensation
Competitive base salary
Visa sponsorship
+2
Principal ML Engineer - Large-Scale Training Performance
Principal ML Engineer - Large-Scale Training Performance

Advanced Micro Devices, Inc. • San Jose (CA)

On-site
USD 130,000 - 160,000