ML Performance Engineer: Scale GPU-Driven Training

Decisive Point

Sunnyvale (CA)

On-site

USD 180,000 - 240,000

Full time

7 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Applied Intuition, Inc. is a Silicon Valley leader powering the future of physical AI. We seek a Performance Engineer to optimize large-scale ML workloads, focusing on distributed training, batch inference, and cost-effective data processing.

You will own profiling across the stack, identify where compute and wall-clock time are spent, and drive improvements that speed time-to-result and reduce processing costs. Join a team that values collaboration and technical excellence.

Qualifications

  • Hands-on ML performance engineering experience with profiling and roofline analysis.
  • Experience with distributed multi-node training at scale and diagnosing scaling inefficiency.
  • Deep familiarity with GPU/accelerator performance concepts including memory bandwidth and kernel launch overhead.
  • Experience with high-throughput or batch inference systems (Triton, TensorRT, ONNX Runtime, Ray).
  • Fluency in Python and proficiency in C++ or another systems language.
  • Excellent debugging, analytical and problem-solving skills.
  • A deep understanding of machine learning foundations and the ability to develop solutions without a playbook.

Responsibilities

  • Profile and optimize distributed training end to end including data loading, preprocessing, and gradient communication.
  • Optimize large-scale offline and batch inference over petabyte-scale logs.
  • Establish roofline and performance models and prioritize optimization opportunities.
  • Improve multi-node scaling efficiency with sharding, parallelism, and communication strategies.
  • Drive cluster goodput by reducing GPU idle time due to data pipelines and I/O bottlenecks.
  • Build benchmarking, observability, and regression-detection tooling for evolving models.
  • Collaborate with engineers across functions to solve large-scale data and compute problems.

Skills

ML performance engineering
Distributed multi-node training
GPU/accelerator concepts
Python
C++
Debugging
Problem solving
ML foundations

Tools

NVIDIA Triton Inference Server
TensorRT
ONNX Runtime
Ray
NCCL

Job description

Applied Intuition, Inc. is a Silicon Valley leader powering the future of physical AI. We seek a Performance Engineer to optimize large-scale ML workloads, focusing on distributed training, batch inference, and cost-effective data processing.

You will own profiling across the stack, identify where compute and wall-clock time are spent, and drive improvements that speed time-to-result and reduce processing costs. Join a team that values collaboration and technical excellence.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

ML Performance Engineer — Scale Distributed Training & Throughput
ML Performance Engineer — Scale Distributed Training & Throughput

applied • Sunnyvale (CA)

On-site
USD 170,000 - 230,000
ML Performance Engineer: Scale Training & Throughput
ML Performance Engineer: Scale Training & Throughput

Applied Intuition • Sunnyvale (CA)

On-site
USD 180,000 - 240,000
Senior ML Performance Engineer: Scale & Throughput
Senior ML Performance Engineer: Scale & Throughput

NLP PEOPLE • Sunnyvale (CA)

On-site
USD 215,000 - 285,000
Staff ML Performance Engineer: Scale Training Throughput
Staff ML Performance Engineer: Scale Training Throughput

Wayve • Sunnyvale (CA)

On-site
USD 130,000 - 160,000
ML Performance Engineer: GPU/CUDA at Scale
ML Performance Engineer: GPU/CUDA at Scale

Selby Jennings • Chicago (IL)

On-site
USD 140,000 - 210,000
Lead GPU ML Training Performance Engineer
Lead GPU ML Training Performance Engineer

Advanced Micro Devices, Inc. • San Jose (CA)

On-site
USD 120,000 - 160,000
Engineering Manager, ML Platform & GPU Infra
Engineering Manager, ML Platform & GPU Infra

applied • Sunnyvale (CA)

On-site
USD 204,000 - 343,000
Equity options
Health insurance
401k retirement benefits
+1
Remote ML Performance Engineer: Optimize Training Inference
Remote ML Performance Engineer: Optimize Training Inference

Bright Vision Technologies • Bellevue (WA)

On-site
USD 100,000 - 150,000
GPU Performance Engineer: Scale ML Inference & Systems
GPU Performance Engineer: Scale ML Inference & Systems

Anthropic • New York (NY)

Hybrid
USD 280,000 - 850,000
Machine Learning Performance Engineer - Offboard Training & Inference
Machine Learning Performance Engineer - Offboard Training & Inference

Decisive Point • Sunnyvale (CA)

On-site
USD 180,000 - 240,000