ML Performance Engineer: Scale Training & Throughput

Applied Intuition

Sunnyvale (CA)

On-site

USD 180,000 - 240,000

Full time

8 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Applied Intuition in Sunnyvale is seeking a Performance Engineer to accelerate large-scale ML workloads in data centers. You will optimize distributed training across many nodes and improve batch inference on petabyte-scale sensor logs, targeting throughput and cost-per-data processed.

You will own profiling across the stack, from data loading to kernel execution, and work with accelerators, ML frameworks, and data infrastructure to close performance gaps and drive faster iterations.

Qualifications

  • Hands-on ML performance engineering experience: profiling, roofline analysis, throughput optimization, root-cause investigation.
  • Experience with distributed multi-node training at scale (FSDP, DeepSpeed, Megatron, NCCL, or equivalent).
  • Deep familiarity with GPU or accelerator performance concepts - memory bandwidth, kernel launch overhead, occupancy, quantization, collective communication.
  • Experience with high-throughput or batch inference systems (NVIDIA Triton Inference Server, TensorRT, ONNX Runtime, Ray, or similar).
  • Fluency in Python and proficiency in C++ or another systems language.
  • Excellent debugging, analytical, and problem-solving skills.
  • Deep understanding of machine learning foundations and ability to develop solutions for problems without established playbooks.

Responsibilities

  • Profile and optimize distributed training end to end across data loading, preprocessing, kernel execution, and checkpointing.
  • Optimize large-scale offline and batch inference over petabyte-scale sensor logs.
  • Establish roofline and performance models and rank optimization opportunities by impact.
  • Improve multi-node scaling efficiency through sharding, parallelism, and memory/bandwidth optimization.
  • Drive cluster goodput by reducing GPU idle time due to I/O, scheduling gaps, or stragglers.
  • Build benchmarking, observability, and regression-detection tooling to protect performance as models evolve.
  • Collaborate with engineers across functions to solve data and compute problems at scale.
  • Contribute to a culture of collaboration, technical excellence, and innovation.

Skills

Profiling & roofline analysis
Distributed multi-node training
Python
C++
GPU / accelerator knowledge
Troubleshooting & debugging

Tools

NVIDIA Triton Inference Server
TensorRT
ONNX Runtime
Kubernetes
Ray

Job description

Applied Intuition in Sunnyvale is seeking a Performance Engineer to accelerate large-scale ML workloads in data centers. You will optimize distributed training across many nodes and improve batch inference on petabyte-scale sensor logs, targeting throughput and cost-per-data processed.

You will own profiling across the stack, from data loading to kernel execution, and work with accelerators, ML frameworks, and data infrastructure to close performance gaps and drive faster iterations.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior ML Performance Engineer: Scale & Throughput
Senior ML Performance Engineer: Scale & Throughput

NLP PEOPLE • Sunnyvale (CA)

On-site
USD 215,000 - 285,000
ML Performance Engineer — Scale Distributed Training & Throughput
ML Performance Engineer — Scale Distributed Training & Throughput

applied • Sunnyvale (CA)

On-site
USD 170,000 - 230,000
ML Performance Engineer: Scale GPU-Driven Training
ML Performance Engineer: Scale GPU-Driven Training

Decisive Point • Sunnyvale (CA)

On-site
USD 180,000 - 240,000
Staff ML Performance Engineer: Scale Training Throughput
Staff ML Performance Engineer: Scale Training Throughput

Wayve • Sunnyvale (CA)

On-site
USD 130,000 - 160,000
ML Performance Engineering Manager
ML Performance Engineering Manager

Google • United States

On-site
USD 207,000 - 300,000
Health, dental, vision
401(k) with company match
Paid time off 20 days
+4
ML Performance Engineer: GPU/CUDA at Scale
ML Performance Engineer: GPU/CUDA at Scale

Selby Jennings • Chicago (IL)

On-site
USD 140,000 - 210,000
Systems ML Engineer: Cloud & Edge Performance
Systems ML Engineer: Cloud & Edge Performance

S27a • Cambridge (MA)

On-site
USD 170,000 - 240,000
Competitive compensation
Full benefits
Production ML Engineer — Scale & Train LLMs
Production ML Engineer — Scale & Train LLMs

Anthropic • San Francisco (CA)

On-site
USD 350,000 - 850,000
Generous vacation and parental leave
Flexible working hours
Office space for collaboration
Staff/Sr. ML Compute Efficiency Engineer
Staff/Sr. ML Compute Efficiency Engineer

Socket.dev • Santa Clara (CA)

On-site
USD 130,000 - 180,000
Remote ML Performance Engineer: Optimize Training Inference
Remote ML Performance Engineer: Optimize Training Inference

Bright Vision Technologies • Bellevue (WA)

On-site
USD 100,000 - 150,000