ML Performance Engineer — Scale Distributed Training & Throughput

applied

Sunnyvale (CA)

On-site

USD 170,000 - 230,000

Full time

8 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Applied Intuition, Inc. is seeking a performance engineer to accelerate large-scale ML workloads in the data center. You will own profiling, optimization, and cost-efficiency for distributed training and large offline inferences.

You will work across accelerators, ML frameworks, and data infra, partnering with teams to land improvements that shorten time-to-result and reduce data processing costs. Collaboration and technical excellence are valued.

Qualifications

  • Hands-on ML performance engineering experience: profiling, roofline analysis, throughput optimization, and root-cause investigation in production systems.
  • Experience with distributed multi-node training at scale (FSDP, DeepSpeed, Megatron, NCCL, or equivalent).
  • Deep familiarity with GPU/accelerator performance concepts: memory bandwidth, kernel launch overhead, occupancy, quantization, collective communication.
  • Experience with high-throughput or batch inference systems: Triton/TensorRT/ONNX etc.

Responsibilities

  • Profile and optimize distributed training end-to-end: data loading, preprocessing, augmentation, kernel execution, gradient communication, checkpointing.
  • Optimize large-scale offline and batch inference over petabyte-scale logs: batching, scheduling, quantization, graph optimization, accelerator saturation.
  • Establish roofline and performance models, quantify gaps between achieved and theoretical performance, rank opportunities by impact.
  • Collaborate across teams to improve multi-node scaling, memory bandwidth, and scheduling gaps on long-running jobs.
  • Build benchmarking, observability, and regression-detection tooling to prevent performance degradation as models evolve.

Skills

ML performance engineering
Distributed training
GPU performance
Python
C++
Debugging

Education

Bachelor's degree in CS/Math/related

Tools

NVIDIA Triton
TensorRT
ONNX Runtime
DeepSpeed

Job description

Applied Intuition, Inc. is seeking a performance engineer to accelerate large-scale ML workloads in the data center. You will own profiling, optimization, and cost-efficiency for distributed training and large offline inferences.

You will work across accelerators, ML frameworks, and data infra, partnering with teams to land improvements that shorten time-to-result and reduce data processing costs. Collaboration and technical excellence are valued.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

ML Performance Engineer: Scale Training & Throughput
ML Performance Engineer: Scale Training & Throughput

Applied Intuition • Sunnyvale (CA)

On-site
USD 180,000 - 240,000
ML Performance Engineer: Scale GPU-Driven Training
ML Performance Engineer: Scale GPU-Driven Training

Decisive Point • Sunnyvale (CA)

On-site
USD 180,000 - 240,000
Senior ML Performance Engineer: Scale & Throughput
Senior ML Performance Engineer: Scale & Throughput

NLP PEOPLE • Sunnyvale (CA)

On-site
USD 215,000 - 285,000
Staff ML Performance Engineer: Scale Training Throughput
Staff ML Performance Engineer: Scale Training Throughput

Wayve • Sunnyvale (CA)

On-site
USD 130,000 - 160,000
ML Performance Engineer: GPU/CUDA at Scale
ML Performance Engineer: GPU/CUDA at Scale

Selby Jennings • Chicago (IL)

On-site
USD 140,000 - 210,000
Staff/Sr. ML Compute Efficiency Engineer
Staff/Sr. ML Compute Efficiency Engineer

Socket.dev • Santa Clara (CA)

On-site
USD 130,000 - 180,000
Performance Engineer, Large-Scale ML Systems
Performance Engineer, Large-Scale ML Systems

Anthropic • New York (NY)

Hybrid
USD 280,000 - 850,000
Competitive compensation
Optional equity donation matching
Generous vacation and parental leave
+1
Remote-Ready ML Systems Engineer: High-Throughput Training
Remote-Ready ML Systems Engineer: High-Throughput Training

United States Digital Space LLC • United States

Hybrid
USD 120,000 - 160,000
ML Performance Engineer — Scalable DL Pipelines & Optimization
ML Performance Engineer — Scalable DL Pipelines & Optimization

Optiver US LLC • New York (NY)

On-site
USD 160,000 - 260,000
Competitive compensation package
Global profit-sharing pool
401(k) match up to 50%
+2
Machine Learning Performance Engineer - Offboard Training & Inference
Machine Learning Performance Engineer - Offboard Training & Inference

Decisive Point • Sunnyvale (CA)

On-site
USD 180,000 - 240,000