Senior ML Infrastructure Engineer - Scalable GPU Training

ATOMS Careers page

San Francisco (CA)

On-site

USD 224,000 - 280,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Medical, Dental, Vision Insurance
401(k)
Unlimited Flexible Time Off

Job summary

ATOMS Careers page is seeking a Staff Machine Learning Infrastructure Engineer to design and implement ML training infrastructure in San Francisco. The role involves building high-performance training pipelines and managing distributed GPU workloads.

Ideal candidates should have 8+ years of software engineering experience, strong backend programming skills, and expertise in Kubernetes. The position offers a competitive salary range of $224,000 - $280,000 per year, along with comprehensive benefits.

Qualifications

  • 8+ years of professional software engineering experience.
  • Proficiency in backend systems programming with relevant languages.
  • Hands-on experience with container orchestration using Kubernetes.

Responsibilities

  • Design and build training infrastructure for ML models.
  • Leverage distributed compute frameworks for ML training jobs.
  • Build data ingestion pipelines for large-scale data processing.

Skills

Software engineering experience
Backend programming skills (Go, Python, Java)
Kubernetes
Distributed ML frameworks (e.g., Ray)
MLOps pipelines
High-throughput data pipelines

Job description

ATOMS Careers page is seeking a Staff Machine Learning Infrastructure Engineer to design and implement ML training infrastructure in San Francisco. The role involves building high-performance training pipelines and managing distributed GPU workloads.

Ideal candidates should have 8+ years of software engineering experience, strong backend programming skills, and expertise in Kubernetes. The position offers a competitive salary range of $224,000 - $280,000 per year, along with comprehensive benefits.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior ML Infrastructure Engineer - GPU Training & MLOps
Senior ML Infrastructure Engineer - GPU Training & MLOps

Atoms • San Francisco (CA)

On-site
USD 224,000 - 280,000
Medical, Dental, Vision, Disability, and Life Insurance
Flexible Spending Account / Health Savings Account Options
401(k)
+2
Senior ML Infrastructure Engineer - Real-World AI Systems
Senior ML Infrastructure Engineer - Real-World AI Systems

Atoms • San Francisco (CA)

On-site
USD 176,000 - 230,000
Medical, dental, and vision insurance
401(k)
Flexible Spending Accounts
+1
Senior ML Training Systems Engineer - Distributed GPU Infra
Senior ML Training Systems Engineer - Distributed GPU Infra

Baseten • San Francisco (CA)

On-site
USD 150,000 - 200,000
Competitive compensation, including equity
100% coverage of medical, dental, and vision insurance
Generous PTO policy
+2
Senior ML Infrastructure Engineer - GPU Training & MLOps
Senior ML Infrastructure Engineer - GPU Training & MLOps

Cssmerge • San Francisco (CA)

On-site
USD 224,000 - 280,000
Medical, Dental, Vision insurance
401(k)
Equity awards
+2
Staff GPU Compute & Infra Automation Engineer (Equity)
Staff GPU Compute & Infra Automation Engineer (Equity)

ATOMS Careers page • San Francisco (CA)

On-site
USD 224,000 - 284,000
Medical, Dental, Vision, Disability, and Life Insurance
Flexible Spending Account / Health Savings Account Options
401(k)
+3
Senior ML Infra Architect — Large-Scale GPU Training
Senior ML Infra Architect — Large-Scale GPU Training

Hark, Inc. • San Jose (CA)

On-site
USD 180,000 - 450,000
Staff GPU Cluster Automation Engineer
Staff GPU Cluster Automation Engineer

Atoms • San Francisco (CA)

On-site
USD 224,000 - 284,000
Medical, Dental, Vision Insurance
401(k)
Unlimited Flexible Time Off
+1
ML Infrastructure Engineer
ML Infrastructure Engineer

Lattice, Inc. • San Francisco (CA)

Hybrid
USD 200,000 - 280,000
Competitive salary
Premium health, dental, and vision insurance
Unlimited PTO
+2
Senior ML Infrastructure Engineer - GPU & Scale
Senior ML Infrastructure Engineer - GPU & Scale

TensorWave • Las Vegas (NV)

On-site
USD 120,000 - 150,000
Competitive Salary
Stock Options
100% paid Medical, Dental, and Vision insurance
+8
Senior ML Infra Engineer: GPU-Optimized Kubernetes Platform
Senior ML Infra Engineer: GPU-Optimized Kubernetes Platform

Hamilton Barnes Associates Limited • San Francisco (CA)

On-site
USD 225,000 - 275,000
Stock options