Senior ML Performance Engineer: Scale & Throughput

NLP PEOPLE

Sunnyvale (CA)

On-site

USD 215,000 - 285,000

Full time

9 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Applied Intuition, Inc. in Sunnyvale, CA, is seeking a performance engineer to optimize large-scale ML workloads in the datacenter.

This role focuses on distributed training across many nodes and high-throughput batch inference over petabytes of real-world autonomy logs for data mining and evaluation. You will own the gap between theoretical capability of accelerators and actual workload performance, performing profiling across the stack and closing the difference to improve training

Qualifications

  • Hands-on ML performance engineering experience with profiling and roofline analysis.
  • Experience with distributed multi-node training at scale and diagnosing scaling inefficiency.
  • Deep familiarity with GPU/accelerator performance concepts and memory bandwidth.
  • Experience with high-throughput or batch inference systems such as Triton or TensorRT.
  • Fluency in Python; proficiency in C++ or another systems language.
  • Strong debugging, analytical, and problem-solving skills.

Responsibilities

  • Profile and optimize distributed training end-to-end across data loading, preprocessing, and communication.
  • Optimize large-scale offline and batch inference over petabyte-scale sensor logs.
  • Establish roofline and performance models; rank optimization opportunities by impact and effort.
  • Improve multi-node scaling efficiency and manage interconnect utilization.
  • Drive cluster goodput by reducing GPU idle time in pipelines and checkpoints.
  • Build benchmarking, observability, and regression-detection tooling.
  • Collaborate with engineers across functions to solve data and compute problems at scale.

Skills

Performance engineering
Distributed training
Python
C++
GPU/accelerator concepts
Profiling tools
Multi-node scaling
Problem solving
ML fundamentals

Tools

NVIDIA Triton
TensorRT
ONNX Runtime
Ray
NCCL
Kubernetes
Nsight Systems/Compute
PyTorch

Job description

Applied Intuition, Inc. in Sunnyvale, CA, is seeking a performance engineer to optimize large-scale ML workloads in the datacenter.

This role focuses on distributed training across many nodes and high-throughput batch inference over petabytes of real-world autonomy logs for data mining and evaluation. You will own the gap between theoretical capability of accelerators and actual workload performance, performing profiling across the stack and closing the difference to improve training

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

ML Performance Engineer: Scale Training & Throughput
ML Performance Engineer: Scale Training & Throughput

Applied Intuition • Sunnyvale (CA)

On-site
USD 180,000 - 240,000
ML Performance Engineer — Scale Distributed Training & Throughput
ML Performance Engineer — Scale Distributed Training & Throughput

applied • Sunnyvale (CA)

On-site
USD 170,000 - 230,000
ML Performance Engineer: Scale GPU-Driven Training
ML Performance Engineer: Scale GPU-Driven Training

Decisive Point • Sunnyvale (CA)

On-site
USD 180,000 - 240,000
Staff ML Performance Engineer: Scale Training Throughput
Staff ML Performance Engineer: Scale Training Throughput

Wayve • Sunnyvale (CA)

On-site
USD 130,000 - 160,000
Remote ML Performance Engineer: Optimize Training Inference
Remote ML Performance Engineer: Optimize Training Inference

Bright Vision Technologies • Bellevue (WA)

On-site
USD 100,000 - 150,000
Systems ML Engineer: Cloud & Edge Performance
Systems ML Engineer: Cloud & Edge Performance

S27a • Cambridge (MA)

On-site
USD 170,000 - 240,000
Competitive compensation
Full benefits
Senior ML Performance Engineer
Senior ML Performance Engineer

well-funded deeptech startup • California (MO)

On-site
USD 200,000 - 250,000
Staff/Sr. ML Compute Efficiency Engineer
Staff/Sr. ML Compute Efficiency Engineer

Socket.dev • Santa Clara (CA)

On-site
USD 130,000 - 180,000
Senior ML Performance Engineer — Ultra-Fast Inference + Equity
Senior ML Performance Engineer — Ultra-Fast Inference + Equity

well-funded deeptech startup • California (MO)

On-site
USD 200,000 - 250,000
ML Performance Engineer: GPU/CUDA at Scale
ML Performance Engineer: GPU/CUDA at Scale

Selby Jennings • Chicago (IL)

On-site
USD 140,000 - 210,000