Senior ML Performance Engineer: Scale & Throughput

NLP PEOPLE

Sunnyvale (CA)

On-site

USD 215,000 - 285,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Applied Intuition, Inc. in Sunnyvale, CA, is seeking a performance engineer to optimize large-scale ML workloads in the datacenter.

This role focuses on distributed training across many nodes and high-throughput batch inference over petabytes of real-world autonomy logs for data mining and evaluation. You will own the gap between theoretical capability of accelerators and actual workload performance, performing profiling across the stack and closing the difference to improve training

Qualifications

  • Hands-on ML performance engineering experience with profiling and roofline analysis.
  • Experience with distributed multi-node training at scale and diagnosing scaling inefficiency.
  • Deep familiarity with GPU/accelerator performance concepts and memory bandwidth.
  • Experience with high-throughput or batch inference systems such as Triton or TensorRT.
  • Fluency in Python; proficiency in C++ or another systems language.
  • Strong debugging, analytical, and problem-solving skills.

Responsibilities

  • Profile and optimize distributed training end-to-end across data loading, preprocessing, and communication.
  • Optimize large-scale offline and batch inference over petabyte-scale sensor logs.
  • Establish roofline and performance models; rank optimization opportunities by impact and effort.
  • Improve multi-node scaling efficiency and manage interconnect utilization.
  • Drive cluster goodput by reducing GPU idle time in pipelines and checkpoints.
  • Build benchmarking, observability, and regression-detection tooling.
  • Collaborate with engineers across functions to solve data and compute problems at scale.

Skills

Performance engineering
Distributed training
Python
C++
GPU/accelerator concepts
Profiling tools
Multi-node scaling
Problem solving
ML fundamentals

Tools

NVIDIA Triton
TensorRT
ONNX Runtime
Ray
NCCL
Kubernetes
Nsight Systems/Compute
PyTorch

Job description

Applied Intuition, Inc. in Sunnyvale, CA, is seeking a performance engineer to optimize large-scale ML workloads in the datacenter.

This role focuses on distributed training across many nodes and high-throughput batch inference over petabytes of real-world autonomy logs for data mining and evaluation. You will own the gap between theoretical capability of accelerators and actual workload performance, performing profiling across the stack and closing the difference to improve training

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI Performance Engineer: Scale ML Throughput & Efficiency
AI Performance Engineer: Scale ML Throughput & Efficiency

Decisive Point • Sunnyvale (CA)

Hybrid
USD 180,000 - 240,000
AI Performance Engineer: Scale ML Throughput & Training
AI Performance Engineer: Scale ML Throughput & Training

InvestedintheMission • Sunnyvale (CA)

On-site
USD 180,000 - 260,000
AI Performance Engineer for Large-Scale Throughput
AI Performance Engineer for Large-Scale Throughput

Applied Intuition • Sunnyvale (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Staff ML Performance Engineer: Scale Training Throughput
Staff ML Performance Engineer: Scale Training Throughput

Wayve • Sunnyvale (CA)

On-site
USD 130,000 - 160,000
Remote ML Performance Engineer: Throughput & Tuning Expert
Remote ML Performance Engineer: Throughput & Tuning Expert

Bright Vision Technologies • Plymouth (MN)

Remote
USD 100,000 - 150,000
Remote AI Performance Engineer - Scale ML Inference
Remote AI Performance Engineer - Scale ML Inference

United States Digital Space LLC • United States

Remote
USD 75,000 - 100,000
ML Infrastructure Engineer: Scale & Performance
ML Infrastructure Engineer: Scale & Performance

Physical Intelligence • San Francisco (CA)

On-site
USD 150,000 - 230,000
Senior ML Performance Engineer
Senior ML Performance Engineer

well-funded deeptech startup • California (MO)

On-site
USD 200,000 - 250,000
Remote ML Systems Engineer: High-Performance Inference & Scale
Remote ML Systems Engineer: High-Performance Inference & Scale

Bright Vision Technologies • United States

Remote
USD 145,000 - 165,000
Staff ML Performance Engineer: Scale Training & GPU
Staff ML Performance Engineer: Scale Training & GPU

OpenDigital Limited • Sunnyvale (CA), Northern (KY)

Hybrid
USD 336,000 - 359,000