Senior ML Infrastructure Engineer - GPU & Scale

TensorWave

Las Vegas (NV)

On-site

USD 120,000 - 150,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Competitive Salary
Stock Options
100% paid Medical, Dental, and Vision insurance
Flexible PTO
Paid Holidays
401(k)
Parental Leave
Flexible Spending Account
Short Term Disability Insurance
Life and Voluntary Supplemental Insurance
Mental Health Benefits

Job summary

A cutting-edge technology company is seeking a Senior Machine Learning Engineer to build and operate systems that power large-scale machine learning training. This role includes designing ML infrastructure, optimizing performance, and enhancing developer experiences. Candidates should have a Bachelor’s degree in Computer Science and expertise in SLURM, Kubernetes, and GPU workloads. The company offers competitive salary, stock options, full medical benefits, flexible PTO, and a mission-driven culture promoting inclusion.

Qualifications

  • Expertise supporting production ML systems using SLURM and Kubernetes.
  • Strong understanding of GPU-accelerated workloads and distributed systems concepts.
  • Solid Linux fundamentals and experience debugging infrastructure-level issues.
  • Ability to build automation and tooling.

Responsibilities

  • Design and improve ML infrastructure systems supporting distributed workloads.
  • Build workload execution and orchestration patterns across GPU environments.
  • Troubleshoot performance and scalability issues.
  • Partner with teams to improve developer experience.

Skills

Production ML systems
Performance optimization
Cluster operations
Workload orchestration
Automation tooling

Education

Bachelor of Science in Computer Science or related field

Tools

SLURM
Kubernetes
Python
Go

Job description

A cutting-edge technology company is seeking a Senior Machine Learning Engineer to build and operate systems that power large-scale machine learning training. This role includes designing ML infrastructure, optimizing performance, and enhancing developer experiences. Candidates should have a Bachelor’s degree in Computer Science and expertise in SLURM, Kubernetes, and GPU workloads. The company offers competitive salary, stock options, full medical benefits, flexible PTO, and a mission-driven culture promoting inclusion.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior ML Training Systems Engineer - Distributed GPU Infra
Senior ML Training Systems Engineer - Distributed GPU Infra

Baseten • San Francisco (CA)

On-site
USD 150,000 - 200,000
Competitive compensation, including equity
100% coverage of medical, dental, and vision insurance
Generous PTO policy
+2
Senior ML Infra Engineer: Build Scalable LLM Systems
Senior ML Infra Engineer: Build Scalable LLM Systems

ServiceNow • Mountain View (CA)

On-site
USD 130,000 - 180,000
Senior Systems Engineer: ML Pipelines & GPU-Scale Infra
Senior Systems Engineer: ML Pipelines & GPU-Scale Infra

Acceler8 Talent • San Francisco (CA)

On-site
USD 150,000 - 250,000
Competitive salary up to $250k
Opportunity to work with top AI video labs
Small, high-impact team with significant growth trajectory
Senior ML Infrastructure Engineer - GPU Training & MLOps
Senior ML Infrastructure Engineer - GPU Training & MLOps

Atoms • San Francisco (CA)

On-site
USD 224,000 - 280,000
Medical, Dental, Vision, Disability, and Life Insurance
Flexible Spending Account / Health Savings Account Options
401(k)
+2
Senior ML Engineer - Build Scalable ML Pipelines (Remote)
Senior ML Engineer - Build Scalable ML Pipelines (Remote)

thatgamecompany • United States

Remote
USD 120,000 - 195,000
Paid Time Off, Holidays, and Winter Break
Medical, dental, and vision coverage
Pet Insurance
+5
Senior GPU ML Infra Engineer — Mid-Training & Inference
Senior GPU ML Infra Engineer — Mid-Training & Inference

Reflection AI • San Francisco (CA)

On-site
USD 120,000 - 160,000
Top-tier compensation
Comprehensive medical, dental, vision, life, and disability insurance
Fully paid parental leave for all new parents
+2
Staff ML Performance Engineer: Scale Training Throughput
Staff ML Performance Engineer: Scale Training Throughput

Wayve • Sunnyvale (CA)

On-site
USD 130,000 - 160,000
Senior ML Infra Engineer - Scale GPU Clusters, Remote
Senior ML Infra Engineer - Scale GPU Clusters, Remote

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 320,000 - 500,000
Equity
Medical/Dental/Vision coverage
Unlimited PTO
+1
Senior ML & Distributed Systems Engineer
Senior ML & Distributed Systems Engineer

Amazon • Mountain View (CA)

On-site
USD 165,200 - 223,600
Health insurance
401(k) matching
Paid time off
Senior ML Platform Engineer — Build Scalable ML
Senior ML Platform Engineer — Build Scalable ML

SoFi • San Francisco (CA)

On-site
USD 128,000 - 240,000
Comprehensive benefits
Career growth opportunities