AI Performance Engineer: Scale ML Throughput & Training

InvestedintheMission

Sunnyvale (CA)

On-site

USD 180,000 - 260,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Applied Intuition, Inc. is seeking a performance engineer to optimize large-scale ML workloads in the datacenter.

You will own the gap between theoretical accelerator performance and real workloads, profiling across the stack to reduce wasted GPU-hours and improve throughput for training and offline inference. You will work at the intersection of accelerators, ML frameworks, and large-scale data infrastructure, partnering with teams to land performance wins that reduce time-to-result and cost

Qualifications

  • Hands-on ML performance engineering experience with profiling and throughput optimization.
  • Experience with distributed multi-node training at scale (e.g., FSDP, DeepSpeed, Megatron, NCCL).
  • Deep familiarity with GPU/accelerator performance concepts: memory bandwidth, kernel overhead, occupancy.
  • Experience with high-throughput or batch inference systems (e.g., NVIDIA Triton, TensorRT, ONNX Runtime).
  • Fluency in Python and proficiency in C++ or another systems language.
  • Excellent debugging, analytical, and problem-solving skills.
  • Strong understanding of ML foundations and ability to solve non-playbook problems.

Responsibilities

  • Profile and optimize distributed training end-to-end: data loading, preprocessing, kernel execution, gradient communication, checkpointing.
  • Optimize large-scale offline and batch inference over petabyte-scale logs: batching, quantization, scheduling, saturation.
  • Establish roofline and performance models; rank optimization opportunities by impact and effort.
  • Improve multi-node scaling efficiency: sharding, parallelism, interconnect utilization, memory bandwidth, kernel fusion.
  • Drive cluster goodput by reducing GPU idle time from input pipelines, I/O, scheduling gaps, and retries.
  • Build benchmarking, observability, and regression-detection tooling to prevent silent degradation.
  • Collaborate with engineers across functions to solve complex data and compute problems at scale.
  • Contribute to a culture of collaboration, technical excellence, and innovation.

Skills

ML performance engineering
Distributed multi-node training
GPU/accelerator concepts
Batch inference systems
Python & C++ proficiency
Debugging & problem solving
ML foundations knowledge

Job description

Applied Intuition, Inc. is seeking a performance engineer to optimize large-scale ML workloads in the datacenter.

You will own the gap between theoretical accelerator performance and real workloads, profiling across the stack to reduce wasted GPU-hours and improve throughput for training and offline inference. You will work at the intersection of accelerators, ML frameworks, and large-scale data infrastructure, partnering with teams to land performance wins that reduce time-to-result and cost

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Performance Engineer: Scale ML Training & Inference
AI Performance Engineer: Scale ML Training & Inference

applied • Sunnyvale (CA)

On-site
USD 180,000 - 240,000
AI Performance Engineer: Scale ML Throughput & Efficiency
AI Performance Engineer: Scale ML Throughput & Efficiency

Decisive Point • Sunnyvale (CA)

Hybrid
USD 180,000 - 240,000
AI Performance Engineer for Large-Scale Throughput
AI Performance Engineer for Large-Scale Throughput

Applied Intuition • Sunnyvale (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Senior ML Performance Engineer: Scale & Throughput
Senior ML Performance Engineer: Scale & Throughput

NLP PEOPLE • Sunnyvale (CA)

On-site
USD 215,000 - 285,000
Staff ML Performance Engineer: Scale Training Throughput
Staff ML Performance Engineer: Scale Training Throughput

Wayve • Sunnyvale (CA)

On-site
USD 130,000 - 160,000
GPU Performance Engineer: Scale ML Inference & Systems
GPU Performance Engineer: Scale ML Inference & Systems

Anthropic • New York (NY)

Hybrid
USD 280,000 - 850,000
AI Performance Engineer
AI Performance Engineer

applied • Sunnyvale (CA)

On-site
USD 180,000 - 240,000
Machine Learning Performance Engineer – Offboard Training & Inference
Machine Learning Performance Engineer – Offboard Training & Inference

NLP PEOPLE • Sunnyvale (CA)

On-site
USD 215,000 - 285,000
ML Inference Performance Visibility Engineer
ML Inference Performance Visibility Engineer

Etched • San Jose (CA)

On-site
USD 150,000 - 210,000
Housing subsidy
Relocation support
Medical benefits
+1
Lead GPU ML Training Performance Engineer
Lead GPU ML Training Performance Engineer

Advanced Micro Devices, Inc. • San Jose (CA)

On-site
USD 120,000 - 160,000