Senior ML Performance Engineer - Distributed Training

Odyssey

Santa Clara (CA)

On-site

USD 120,000 - 160,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

An AI lab in Santa Clara is seeking a skilled software engineer with over 8 years of experience to optimize machine learning models for real-time applications. The role involves designing distributed training strategies, collaborating with ML researchers, and developing tools for performance enhancement. Ideal candidates have a deep understanding of modern machine learning architectures and proficiency with PyTorch and NVIDIA GPU ecosystems.

Qualifications

  • 8+ years of software engineering experience, with significant work in ML performance.
  • Deep insight into modern machine learning architectures.
  • Proficiency with PyTorch and optimization stacks.

Responsibilities

  • Optimize models for real-time use by hundreds of thousands of users.
  • Design and implement distributed training strategies.
  • Develop tools to identify performance bottlenecks.

Skills

Software engineering experience
Machine learning performance
Problem-solving skills
PyTorch proficiency
Deep understanding of ML architectures

Tools

NVIDIA GPU ecosystems
Triton

Job description

An AI lab in Santa Clara is seeking a skilled software engineer with over 8 years of experience to optimize machine learning models for real-time applications. The role involves designing distributed training strategies, collaborating with ML researchers, and developing tools for performance enhancement. Ideal candidates have a deep understanding of modern machine learning architectures and proficiency with PyTorch and NVIDIA GPU ecosystems.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

ML Performance Engineer – Real-Time Inference
ML Performance Engineer – Real-Time Inference

Odyssey • Palo Alto (CA)

On-site
USD 130,000 - 160,000
Senior ML Training Systems Engineer - Distributed GPU Infra
Senior ML Training Systems Engineer - Distributed GPU Infra

Baseten • San Francisco (CA)

On-site
USD 150,000 - 200,000
Competitive compensation, including equity
100% coverage of medical, dental, and vision insurance
Generous PTO policy
+2
Staff ML Performance Engineer: Scale Training Throughput
Staff ML Performance Engineer: Scale Training Throughput

Wayve • Sunnyvale (CA)

On-site
USD 130,000 - 160,000
Staff ML Systems Engineer — Distributed Training at Scale
Staff ML Systems Engineer — Distributed Training at Scale

RadixArk • Palo Alto (CA)

On-site
USD 120,000 - 160,000
Comprehensive benefits
Flexible work arrangements
Staff Engineer - Training & Inference for Distributed AI
Staff Engineer - Training & Inference for Distributed AI

Boson AI • Santa Clara (CA)

On-site
USD 150,000 - 600,000
Senior ML Performance Engineer: Scale & Throughput
Senior ML Performance Engineer: Scale & Throughput

NLP PEOPLE • Sunnyvale (CA)

On-site
USD 215,000 - 285,000
ML Performance Engineer — Scalable DL Pipelines & Optimization
ML Performance Engineer — Scalable DL Pipelines & Optimization

Optiver US LLC • New York (NY)

On-site
USD 160,000 - 260,000
Competitive compensation package
Global profit-sharing pool
401(k) match up to 50%
+2
Research Software Engineer — Scalable RL & Distributed Training
Research Software Engineer — Scalable RL & Distributed Training

Reflection AI • San Francisco (CA)

On-site
USD 120,000 - 160,000
Top-tier compensation
Comprehensive health insurance
Fully paid parental leave
+2
Senior AI/ML Distributed Training Engineer
Senior AI/ML Distributed Training Engineer

Amazon • Cupertino (CA)

On-site
USD 193,300 - 261,500
Senior AI Training Performance Architect — Optimize at Scale
Senior AI Training Performance Architect — Optimize at Scale

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 210,000 - 340,000
Equity
Benefits