Lead ML Training Optimizer for Scalable AI Systems

3M HEALTHCARE

San Francisco, Phoenix, Pittsburgh, Dallas (CA, AZ, Allegheny County, TX)

Hybrid

USD 140,000 - 210,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Competitive compensation

Job summary

Waabi is seeking a highly skilled engineer to develop distributed training frameworks and optimize performance for research and production environments. You will profile runtime/memory, identify new technologies such as CUDA kernels and quantization, and collaborate with researchers to implement best practices for efficient resource usage.

You will also design and improve tooling and dashboards to promote broad adoption across teams, contributing to Waabi’s cutting-edge autonomous driving

Qualifications

  • MS/PhD or Bachelor's with 4+ years in CS/Robotics or related field
  • Strong coding skills in Python, C++, or Rust
  • Experience with DL frameworks like PyTorch or JAX
  • Experience profiling CPU/GPU using profiling tools
  • Collaborative team player with passion for self-driving tech

Responsibilities

  • Build standardized distributed training frameworks for research and production.
  • Profile model runtime and memory to optimize performance.
  • Evaluate emerging technologies for Waabi’s training/inference frameworks (CUDA kernels, quantization, deployment).
  • Collaborate with researchers and ML engineers on best practices for resource usage.
  • Create tooling and dashboards to drive broad adoption of work.

Skills

Python
C++
Rust
PyTorch
Nsight

Education

MS/PhD or Bachelor's with 4+ years experience

Tools

NVIDIA Nsight
PyTorch Profiler
CUDA
Bazel
Kubernetes

Job description

Waabi is seeking a highly skilled engineer to develop distributed training frameworks and optimize performance for research and production environments. You will profile runtime/memory, identify new technologies such as CUDA kernels and quantization, and collaborate with researchers to implement best practices for efficient resource usage.

You will also design and improve tooling and dashboards to promote broad adoption across teams, contributing to Waabi’s cutting-edge autonomous driving

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior / Staff ML Training Optimization Engineer
Senior / Staff ML Training Optimization Engineer

3M HEALTHCARE • San Francisco (CA), Phoenix (AZ), Pittsburgh, Dallas (TX)

Hybrid
USD 140,000 - 210,000
Competitive compensation
Senior ML Training Architect — Distributed, Flexible & Remote
Senior ML Training Architect — Distributed, Flexible & Remote

Waabi • United States

Hybrid
USD 141,000 - 249,000
Competitive compensation and equity awards
Health and Wellness benefits (Medical, Dental, Vision)
Unlimited Vacation
+4
Senior ML Data Pipelines Engineer — Remote & Flexible Hours
Senior ML Data Pipelines Engineer — Remote & Flexible Hours

ProducePay • United States

On-site
USD 148,000 - 260,000
Competitive compensation and equity
Health and Wellness benefits (Medical,
Unlimited Vacation
+3
Senior ML Onboard Optimization Engineer — Remote & Flexible
Senior ML Onboard Optimization Engineer — Remote & Flexible

Waabi • United States

On-site
USD 141,000 - 249,000
Competitive compensation and equity awards
Health and Wellness benefits
Unlimited Vacation
+4
Senior ML Platform & MLOps Engineer
Senior ML Platform & MLOps Engineer

Waabi • United States

Hybrid
USD 157,000 - 234,000
Competitive compensation and equity awards
Health and wellness benefits
Unlimited vacation
+3
Senior / Staff ML Training Optimization Engineer
Senior / Staff ML Training Optimization Engineer

Waabi • United States

Hybrid
USD 141,000 - 249,000
Competitive compensation and equity awards
Health and Wellness benefits (Medical, Dental, Vision)
Unlimited Vacation
+4
Senior / Staff Software Engineer, ML Datasets & Data Pipelines
Senior / Staff Software Engineer, ML Datasets & Data Pipelines

ProducePay • United States

On-site
USD 148,000 - 260,000
Competitive compensation and equity
Health and Wellness benefits (Medical,
Unlimited Vacation
+3
Distributed ML Engineer for High-Performance AI Training
Distributed ML Engineer for High-Performance AI Training

Ifm Us • Sunnyvale (CA)

On-site
USD 140,000 - 210,000
Medical benefits
Dental benefits
Vision benefits
+7
Senior ML Systems Engineer — Auto Labelling & Perception
Senior ML Systems Engineer — Auto Labelling & Perception

Waabi • Dallas (TX)

On-site
USD 170,000 - 220,000
Competitive compensation and equity awards
Health and Wellness benefits
Unlimited Vacation
+3
Senior Staff Software Engineer - AV Simulation Platform
Senior Staff Software Engineer - AV Simulation Platform

ProducePay • United States

Hybrid
USD 159,000 - 268,000
Equity awards
Health benefits
Unlimited vacation
+2