AI Performance Engineer: Scale ML Throughput & Efficiency

Decisive Point

Sunnyvale (CA)

Hybrid

USD 180,000 - 240,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Applied Intuition, Inc. is seeking a performance engineer to optimize large-scale machine learning workloads in datacenters. You will focus on distributed training across many nodes and high-throughput inference over vast autonomy logs.

The role centers on throughput, goodput, and cost per data unit, with ownership of profiling and bottleneck resolution. The ideal candidate has hands-on ML performance engineering experience, multi-node training scale familiarity, and strong Python/C++ skills.

Qualifications

  • Hands-on ML performance engineering: profiling, roofline, throughput optimization, root-cause investigation.

Responsibilities

  • Profile and optimize distributed training end to end: data loading, preprocessing, kernel execution, gradient communication, checkpointing.
  • Optimize offline/batch inference on petabyte-scale logs: batching, quantization, execution, and saturation across long-running sweeps.
  • Establish roofline models and quantify gaps between achieved and theoretical performance; rank opportunities by impact and effort.
  • Improve multi-node scaling efficiency: sharding, parallelism, collective communication, memory bandwidth, kernel fusion bottlenecks.
  • Drive cluster goodput by reducing GPU idle time due to data pipelines, I/O, scheduling gaps, stragglers, and recoveries.
  • Build benchmarking, observability, and regression tooling; collaborate across teams to solve scale problems.

Skills

ML performance
Distributed training
GPU performance
Batch inference
Python & C++
Debugging & analytics
ML foundations

Tools

NVIDIA Triton Inference Server
TensorRT
ONNX Runtime
Ray
Nsight Systems/Compute
PyTorch Profiler

Job description

Applied Intuition, Inc. is seeking a performance engineer to optimize large-scale machine learning workloads in datacenters. You will focus on distributed training across many nodes and high-throughput inference over vast autonomy logs.

The role centers on throughput, goodput, and cost per data unit, with ownership of profiling and bottleneck resolution. The ideal candidate has hands-on ML performance engineering experience, multi-node training scale familiarity, and strong Python/C++ skills.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior ML Performance Engineer: Scale & Throughput
Senior ML Performance Engineer: Scale & Throughput

NLP PEOPLE • Sunnyvale (CA)

On-site
USD 215,000 - 285,000
AI Performance Engineer: Scale ML Workloads at Datacenter
AI Performance Engineer: Scale ML Workloads at Datacenter

Applied Intuition • Sunnyvale (CA)

On-site
USD 215,000 - 285,000
Staff ML Performance Engineer: Scale Training Throughput
Staff ML Performance Engineer: Scale Training Throughput

Wayve • Sunnyvale (CA)

On-site
USD 130,000 - 160,000
Remote ML Performance Engineer – Scale AI Throughput
Remote ML Performance Engineer – Scale AI Throughput

Bright Vision Technologies • Bothell (WA)

On-site
USD 100,000 - 150,000
AI Performance Engineer – HPC, ARM & Distributed Inference
AI Performance Engineer – HPC, ARM & Distributed Inference

EngineersOfAI • Austin (TX)

On-site
USD 90,000 - 120,000
Distributed ML Training Performance Engineer
Distributed ML Training Performance Engineer

OpenAI • California (MO)

Hybrid
USD 170,000 - 260,000
Relocation assistance
ML Performance Engineer — Scalable DL Pipelines & Optimization
ML Performance Engineer — Scalable DL Pipelines & Optimization

Optiver US LLC • New York (NY)

On-site
USD 160,000 - 260,000
Competitive compensation package
Global profit-sharing pool
401(k) match up to 50%
+2
AI/HPC Performance Engineer: Scale Large AI Clusters
AI/HPC Performance Engineer: Scale Large AI Clusters

Meta • Menlo Park (CA)

On-site
USD 154,000 - 217,000
Bonus
Equity
Benefits
ML Infrastructure Engineer: Scale & Performance
ML Infrastructure Engineer: Scale & Performance

Physical Intelligence • San Francisco (CA)

On-site
USD 150,000 - 230,000
ML Performance Engineer – Real-Time Inference
ML Performance Engineer – Real-Time Inference

Odyssey • Palo Alto (CA)

On-site
USD 130,000 - 160,000