AI Performance Engineer: Scale ML Workloads at Datacenter

Applied Intuition

Sunnyvale (CA)

On-site

USD 215,000 - 285,000

Full time

7 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Applied Intuition in Sunnyvale, CA is seeking a performance engineer to optimize large-scale ML workloads across datacenters. You will own profiling, roofline analysis, and end-to-end throughput improvements for training and offline inference on petabyte-scale logs.

Collaborate with teams across accelerators, ML frameworks, and data infrastructure; drive models of performance, reduce GPU idle, and build tooling for observability and regression detection.

Qualifications

  • Hands-on ML performance engineering experience: profiling, roofline analysis, throughput optimization, and root-cause investigation in production systems.
  • Experience with distributed multi-node training at scale (FSDP, DeepSpeed, Megatron, NCCL, or equivalent), including diagnosing scaling inefficiency as node count grows.
  • Deep familiarity with GPU or accelerator performance concepts - memory bandwidth, kernel launch overhead, occupancy, quantization, collective communication.
  • Experience with high-throughput or batch inference systems (NVIDIA Triton Inference Server, TensorRT, ONNX Runtime, Ray, or similar).

Responsibilities

  • Profile and optimize distributed training end to end - data loading and preprocessing, augmentation, kernel execution, gradient communication, and checkpointing
  • Optimize large-scale offline and batch inference over petabyte-scale sensor logs: batching and scheduling strategies, quantization and low-precision execution, graph optimization, and accelerator saturation across long-running sweeps
  • Establish roofline and performance models for our workloads, quantify the gap between achieved and theoretical performance, and stack-rank optimization opportunities by impact and effort
  • Improve multi-node scaling efficiency: sharding and parallelism strategies, collective communication, interconnect utilization, and memory-bandwidth and kernel-fusion bottlenecks
  • Drive cluster goodput - reduce GPU idle time from input pipeline stalls, storage and network I/O, scheduling gaps, stragglers, and failure recovery on long-running jobs
  • Build the benchmarking, observability, and regression-detection tooling that keeps performance from silently degrading as models and code evolve
  • Collaborate with engineers across functions to solve complex data and compute problems at scale
  • Contribute to a team culture that values effective collaboration, technical excellence, and innovation

Skills

Distributed multi-node training
GPU/accelerator performance concepts
High-throughput batch inference
Python
C++
Debugging & problem solving
ML fundamentals

Tools

NVIDIA Triton Inference Server
TensorRT
ONNX Runtime
Ray
NCCL
DeepSpeed
Megatron
FSDP

Job description

Applied Intuition in Sunnyvale, CA is seeking a performance engineer to optimize large-scale ML workloads across datacenters. You will own profiling, roofline analysis, and end-to-end throughput improvements for training and offline inference on petabyte-scale logs.

Collaborate with teams across accelerators, ML frameworks, and data infrastructure; drive models of performance, reduce GPU idle, and build tooling for observability and regression detection.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior ML Performance Engineer: Scale & Throughput
Senior ML Performance Engineer: Scale & Throughput

NLP PEOPLE • Sunnyvale (CA)

On-site
USD 215,000 - 285,000
AI Performance Engineer: Scale ML Throughput & Efficiency
AI Performance Engineer: Scale ML Throughput & Efficiency

Decisive Point • Sunnyvale (CA)

Hybrid
USD 180,000 - 240,000
Staff ML Performance Engineer: Scale Training Throughput
Staff ML Performance Engineer: Scale Training Throughput

Wayve • Sunnyvale (CA)

On-site
USD 130,000 - 160,000
Remote ML Performance Engineer – Scale AI Throughput
Remote ML Performance Engineer – Scale AI Throughput

Bright Vision Technologies • Bothell (WA)

On-site
USD 100,000 - 150,000
Senior ML Infra Engineer – End-to-End Pipelines
Senior ML Infra Engineer – End-to-End Pipelines

Applied Intuition • Sunnyvale (CA)

Hybrid
USD 153,000 - 222,000
Health insurance
Dental insurance
Vision insurance
+6
ML Infrastructure Engineer: Scale & Performance
ML Infrastructure Engineer: Scale & Performance

Physical Intelligence • San Francisco (CA)

On-site
USD 150,000 - 230,000
AI Inference Performance & Scale Engineer
AI Inference Performance & Scale Engineer

AMD • San Jose (CA)

On-site
USD 150,000 - 210,000
Benefits at a glance
Distributed ML Training Performance Engineer
Distributed ML Training Performance Engineer

OpenAI • California (MO)

Hybrid
USD 170,000 - 260,000
Relocation assistance
Systems ML Engineer: Performance, Cloud/Edge, Equity
Systems ML Engineer: Performance, Cloud/Edge, Equity

Breakout Ventures • Cambridge (MA)

On-site
USD 180,000 - 270,000
Equity
Lunch subsidy
Health insurance
+1
Senior AI Systems Performance Engineer
Senior AI Systems Performance Engineer

Goaly • Menlo Park (CA), Northern (KY)

Hybrid
USD 150,000 - 180,000