AI Performance Engineer for Large-Scale Throughput

Applied Intuition

Sunnyvale, Northern (CA, KY)

Hybrid

USD 180,000 - 240,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Applied Intuition, Inc. is building the digital infrastructure for autonomous systems with a focus on large-scale ML workloads.

We are seeking a performance engineer to optimize distributed training and offline inference across multi-node clusters, targeting throughput and cost per unit of data processed. You will own profiling across accelerators, ML frameworks, and data infrastructure, driving improvements that translate to faster time-to-result and lower processing costs.

Qualifications

  • Hands‑on ML performance engineering experience: profiling, roofline analysis, throughput optimization, and root‑cause investigation in production systems.
  • Experience with distributed multi-node training at scale (FSDP, DeepSpeed, Megatron, NCCL, or equivalent), including diagnosing scaling inefficiency as node count grows
  • Deep familiarity with GPU or accelerator performance concepts - memory bandwidth, kernel launch overhead, occupancy, quantization, collective communication
  • Experience with high‑throughput or batch inference systems (NVIDIA Triton Inference Server, TensorRT, ONNX Runtime, Ray, or similar)
  • Fluency in Python and proficiency in C++ or another systems language
  • Excellent debugging, analytical, and problem‑solving skills
  • A deep understanding of machine learning foundations, and the ability to develop technical solutions for problems with no established playbook

Responsibilities

  • Profile and optimize distributed training end to end - data loading and preprocessing, augmentation, kernel execution, gradient communication, and checkpointing
  • Optimize large-scale offline and batch inference over petabyte-scale sensor logs: batching and scheduling strategies, quantization and low-precision execution, graph optimization, and accelerator saturation across long-running sweeps
  • Establish roofline and performance models for our workloads, quantify the gap between achieved and theoretical performance, and stack-rank optimization opportunities by impact and effort
  • Improve multi-node scaling efficiency: sharding and parallelism strategies, collective communication, interconnect utilization, and memory-bandwidth and kernel-fusion bottlenecks
  • Drive cluster goodput - reduce GPU idle time from input pipeline stalls, storage and network I/O, scheduling gaps, stragglers, and failure recovery on long-running jobs
  • Build the benchmarking, observability, and regression-detection tooling that keeps performance from silently degrading as models and code evolve
  • Collaborate with engineers across functions to solve complex data and compute problems at scale
  • Contribute to a team culture that values effective collaboration, technical excellence, and innovation

Skills

ML performance
Distributed training
GPU performance
Inference systems
Python
C++
Debugging
ML foundations

Tools

NVIDIA Triton
TensorRT
ONNX Runtime
Ray

Job description

Applied Intuition, Inc. is building the digital infrastructure for autonomous systems with a focus on large-scale ML workloads.

We are seeking a performance engineer to optimize distributed training and offline inference across multi-node clusters, targeting throughput and cost per unit of data processed. You will own profiling across accelerators, ML frameworks, and data infrastructure, driving improvements that translate to faster time-to-result and lower processing costs.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Performance Engineer: Scale ML Throughput & Efficiency
AI Performance Engineer: Scale ML Throughput & Efficiency

Decisive Point • Sunnyvale (CA)

Hybrid
USD 180,000 - 240,000
AI Performance Engineer: Scale ML Throughput & Training
AI Performance Engineer: Scale ML Throughput & Training

InvestedintheMission • Sunnyvale (CA)

On-site
USD 180,000 - 260,000
AI Performance Engineer: Scale ML Training & Inference
AI Performance Engineer: Scale ML Training & Inference

applied • Sunnyvale (CA)

On-site
USD 180,000 - 240,000
Senior ML Performance Engineer: Scale & Throughput
Senior ML Performance Engineer: Scale & Throughput

NLP PEOPLE • Sunnyvale (CA)

On-site
USD 215,000 - 285,000
AI Performance Engineer
AI Performance Engineer

applied • Sunnyvale (CA)

On-site
USD 180,000 - 240,000
AI/HPC Performance Engineer: Scale Large AI Clusters
AI/HPC Performance Engineer: Scale Large AI Clusters

Meta • Menlo Park (CA)

On-site
USD 154,000 - 217,000
Bonus
Equity
Benefits
Performance Engineer, Inference Engine - High-Performance AI
Performance Engineer, Inference Engine - High-Performance AI

EngineersOfAI • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Machine Learning Performance Engineer – Offboard Training & Inference
Machine Learning Performance Engineer – Offboard Training & Inference

NLP PEOPLE • Sunnyvale (CA)

On-site
USD 215,000 - 285,000
AI Performance Engineer – HPC, ARM & Distributed Inference
AI Performance Engineer – HPC, ARM & Distributed Inference

EngineersOfAI • Austin (TX)

On-site
USD 90,000 - 120,000
Performance Engineer — AI Inference Systems
Performance Engineer — AI Inference Systems

Anthropic • San Francisco (CA)

Hybrid
USD 350,000 - 850,000
Visa sponsorship
Flexible hybrid work policy