AI Performance Engineer: Scale ML Training & Inference

applied

Sunnyvale (CA)

On-site

USD 180,000 - 240,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Applied Intuition, Inc. in Sunnyvale, California, is seeking a performance engineer to optimize large-scale ML workloads in the data center.

You will focus on distributed training across many nodes and high-throughput offline processing of petabyte-scale sensor logs to enable faster iteration. You will profile across the stack, identify bottlenecks, and close the gap between theoretical and actual performance, collaborating with teams across accelerators, ML frameworks, and data infrastructure

Qualifications

  • Hands-on ML performance engineering experience: profiling, roofline analysis, throughput optimization, and root-cause investigation.
  • Experience with distributed multi-node training at scale (FSDP, DeepSpeed, Megatron, NCCL, or equivalent).
  • Deep familiarity with GPU or accelerator performance concepts - memory bandwidth and kernel launch overhead.
  • Experience with high-throughput or batch inference systems (NVIDIA Triton, TensorRT, ONNX Runtime, Ray).
  • Fluency in Python and proficiency in C++ or another systems language.
  • Excellent debugging, analytical, and problem-solving skills.
  • A deep understanding of machine learning foundations and ability to develop technical solutions.

Responsibilities

  • Profile and optimize distributed training end to end across data loading, preprocessing, kernels, and communication.
  • Optimize large-scale offline and batch inference over petabyte-scale sensor logs.
  • Establish roofline models and rank optimization opportunities by impact.
  • Improve multi-node scaling: sharding, communication, interconnect utilization, memory bandwidth.
  • Drive cluster goodput by reducing GPU idle time due to pipeline stalls and I/O gaps.
  • Build benchmarking, observability, and regression-detection tooling.
  • Collaborate with engineers across functions to solve data and compute problems at scale.
  • Contribute to a culture of collaboration, technical excellence, and innovation.

Skills

ML performance
Distributed training
GPU performance
Batch inference
Python & C++
Debugging skills
ML foundations

Tools

NVIDIA Triton
TensorRT
ONNX Runtime
Ray
Nsight Systems/Compute

Job description

Applied Intuition, Inc. in Sunnyvale, California, is seeking a performance engineer to optimize large-scale ML workloads in the data center.

You will focus on distributed training across many nodes and high-throughput offline processing of petabyte-scale sensor logs to enable faster iteration. You will profile across the stack, identify bottlenecks, and close the gap between theoretical and actual performance, collaborating with teams across accelerators, ML frameworks, and data infrastructure

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior ML Performance Engineer: Scale & Throughput
Senior ML Performance Engineer: Scale & Throughput

NLP PEOPLE • Sunnyvale (CA)

On-site
USD 215,000 - 285,000
AI Performance Engineer: Scale ML Throughput & Training
AI Performance Engineer: Scale ML Throughput & Training

InvestedintheMission • Sunnyvale (CA)

On-site
USD 180,000 - 260,000
AI Performance Engineer: Scale ML Throughput & Efficiency
AI Performance Engineer: Scale ML Throughput & Efficiency

Decisive Point • Sunnyvale (CA)

Hybrid
USD 180,000 - 240,000
AI Performance Engineer for Large-Scale Throughput
AI Performance Engineer for Large-Scale Throughput

Applied Intuition • Sunnyvale (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Staff ML Performance Engineer: Scale Training Throughput
Staff ML Performance Engineer: Scale Training Throughput

Wayve • Sunnyvale (CA)

On-site
USD 130,000 - 160,000
ML Infrastructure Engineer: Scale & Performance
ML Infrastructure Engineer: Scale & Performance

Physical Intelligence • San Francisco (CA)

On-site
USD 150,000 - 230,000
ML Inference Performance Visibility Engineer
ML Inference Performance Visibility Engineer

Etched • San Jose (CA)

On-site
USD 150,000 - 210,000
Housing subsidy
Relocation support
Medical benefits
+1
ML Performance Engineer – Real-Time Inference
ML Performance Engineer – Real-Time Inference

Odyssey • Palo Alto (CA)

On-site
USD 130,000 - 160,000
ML Infra Engineer: Scale & Optimize Large-Scale Training
ML Infra Engineer: Scale & Optimize Large-Scale Training

Physical Intelligence • San Francisco (CA)

On-site
USD 180,000 - 240,000
AI Infrastructure Engineer — Scale ML Training & Inference
AI Infrastructure Engineer — Scale ML Training & Inference

Triwill Group • San Francisco (CA), Northern (KY)

Hybrid
USD 140,000 - 190,000