Senior ML Performance Engineer - Distributed Training

Odyssey

Santa Clara (CA)

On-site

USD 120,000 - 160,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

An AI lab in Santa Clara is seeking a skilled software engineer with over 8 years of experience to optimize machine learning models for real-time applications. The role involves designing distributed training strategies, collaborating with ML researchers, and developing tools for performance enhancement. Ideal candidates have a deep understanding of modern machine learning architectures and proficiency with PyTorch and NVIDIA GPU ecosystems.

Qualifications

  • 8+ years of software engineering experience, with significant work in ML performance.
  • Deep insight into modern machine learning architectures.
  • Proficiency with PyTorch and optimization stacks.

Responsibilities

  • Optimize models for real-time use by hundreds of thousands of users.
  • Design and implement distributed training strategies.
  • Develop tools to identify performance bottlenecks.

Skills

Software engineering experience
Machine learning performance
Problem-solving skills
PyTorch proficiency
Deep understanding of ML architectures

Tools

NVIDIA GPU ecosystems
Triton

Job description

An AI lab in Santa Clara is seeking a skilled software engineer with over 8 years of experience to optimize machine learning models for real-time applications. The role involves designing distributed training strategies, collaborating with ML researchers, and developing tools for performance enhancement. Ideal candidates have a deep understanding of modern machine learning architectures and proficiency with PyTorch and NVIDIA GPU ecosystems.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

ML Performance Engineer – Real-Time Inference
ML Performance Engineer – Real-Time Inference

Odyssey • Palo Alto (CA)

On-site
USD 130,000 - 160,000
Principal ML Engineer - Large-Scale Training Performance
Principal ML Engineer - Large-Scale Training Performance

Advanced Micro Devices, Inc. • San Jose (CA)

On-site
USD 130,000 - 160,000
Senior ML Training Systems Engineer - Distributed GPU Infra
Senior ML Training Systems Engineer - Distributed GPU Infra

Baseten • San Francisco (CA)

On-site
USD 150,000 - 200,000
Competitive compensation, including equity
100% coverage of medical, dental, and vision insurance
Generous PTO policy
+2
Principal ML Engineer: Large-Scale Training Performance
Principal ML Engineer: Large-Scale Training Performance

Advanced Micro Devices • San Jose (CA)

Hybrid
USD 120,000 - 180,000
Comprehensive benefits package
Innovative work culture
Staff ML Performance Engineer: Scale Training Throughput
Staff ML Performance Engineer: Scale Training Throughput

Wayve • Sunnyvale (CA)

On-site
USD 130,000 - 160,000
Senior ML Engineer, Distributed Training & P2P Systems
Senior ML Engineer, Distributed Training & P2P Systems

Pluralis Research • California (MO)

Remote
USD 120,000 - 160,000
Principal ML Engineer: Large-Scale Training & Performance
Principal ML Engineer: Large-Scale Training & Performance

AMD • San Jose (CA)

Hybrid
USD 130,000 - 160,000
Staff ML Systems Engineer — Distributed Training at Scale
Staff ML Systems Engineer — Distributed Training at Scale

RadixArk • Palo Alto (CA)

On-site
USD 120,000 - 160,000
Comprehensive benefits
Flexible work arrangements
Staff Engineer - Training & Inference for Distributed AI
Staff Engineer - Training & Inference for Distributed AI

Boson AI • Santa Clara (CA)

On-site
USD 150,000 - 600,000
Senior ML Performance Engineer: Scale & Throughput
Senior ML Performance Engineer: Scale & Throughput

NLP PEOPLE • Sunnyvale (CA)

On-site
USD 215,000 - 285,000