AI Training Performance Engineer — Large-Scale GPU Training

Figure

San Jose (CA)

On-site

USD 200,000 - 400,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Figure, an AI robotics company headquartered in San Jose, CA, is seeking an experienced AI Training Performance Engineer to advance our large-scale model training. You will optimize distributed training across thousands of GPUs, tune GPU kernels, and co-design training recipes with researchers.

You will build tooling and dashboards, improve data pipelines, and extend kernel compilers for non-NVIDIA accelerators.

Qualifications

  • Degree in CS, CE, or related field required.
  • 3+ years in AI performance engineering leading large-scale projects.
  • Expertise in GPU architecture, memory bandwidth, and performance profiling.
  • Strong Python and CUDA/C++ and ability to read/modify framework internals.
  • Experience with NCCL and RDMA/InfiniBand topology-aware placement.

Responsibilities

  • Optimize training performance for 100B+ parameter models across 100k+ GPUs.
  • Collaborate on accelerator choice, cluster topology, scheduling, and hardware procurement.
  • Write and optimize custom kernels (Triton/CUDA).
  • Build tooling and dashboards for performance monitoring and root-cause analysis across training jobs.
  • Optimize data loading and preprocessing pipelines to avoid I/O bottlenecks.
  • Improve checkpointing, fault tolerance, and elastic restart.
  • Co-design model architectures and training recipes for scale.
  • Extend kernel compilers to support non-NVIDIA accelerators.
  • Build agentic systems to benchmark and iterate on custom kernels.
  • Evaluate emerging accelerator architectures and port benchmarks.

Skills

GPU Architecture
Profiling Tools
NCCL & Networking
Python & CUDA/C++
Debugging at scale
Hardware-efficiency metrics

Education

Bachelor's or Master's degree in CS/CE or related field

Tools

Nsight Systems
PyTorch Profiler
HTA

Job description

Figure, an AI robotics company headquartered in San Jose, CA, is seeking an experienced AI Training Performance Engineer to advance our large-scale model training. You will optimize distributed training across thousands of GPUs, tune GPU kernels, and co-design training recipes with researchers.

You will build tooling and dashboards, improve data pipelines, and extend kernel compilers for non-NVIDIA accelerators.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI Training Performance Engineer (GPU & Scale)
Senior AI Training Performance Engineer (GPU & Scale)

figure.ai • San Jose (CA), Northern (KY)

Hybrid
USD 200,000 - 400,000
Senior Training Infra Engineer - 800+ GPU Scale
Senior Training Infra Engineer - 800+ GPU Scale

Figureai • San Jose (CA)

On-site
USD 150,000 - 350,000
Fellow GPU Performance Optimizer for AI Training
Fellow GPU Performance Optimizer for AI Training

Advanced Micro Devices • San Jose (CA)

On-site
USD 140,000 - 180,000
Health insurance
Retirement plan
Paid time off
Senior AI Training Infra Engineer — GPU Clusters
Senior AI Training Infra Engineer — GPU Clusters

Designworks Talent • Bellevue (WA)

Hybrid
USD 180,000 - 300,000
Medical, dental, vision insurance
401(k) with company match
Paid holidays
Senior GPU Performance Engineer for AI Training
Senior GPU Performance Engineer for AI Training

CareerArc • San Jose (CA)

Hybrid
USD 150,000 - 200,000
Competitive salary
Comprehensive benefits
Helix AI Engineer, Training Performance
Helix AI Engineer, Training Performance

figure.ai • San Jose (CA), Northern (KY)

Hybrid
USD 200,000 - 400,000
Remote GPU Performance Engineer: Scale Training & Inference
Remote GPU Performance Engineer: Scale Training & Inference

Reka • United States

Remote
USD 120,000 - 150,000
Five weeks of paid leave
Comprehensive healthcare benefits
Visa support for H1B and OPT transfers
Senior AI Training Performance Architect-Optimization Lead
Senior AI Training Performance Architect-Optimization Lead

Nvidia Corporation • Santa Clara (CA)

On-site
USD 184,000 - 357,000
AI Engineer: Model Training, Inference & GPU Infra
AI Engineer: Model Training, Inference & GPU Infra

Agentrys • San Jose (CA), Northern (KY)

Hybrid
USD 170,000 - 210,000
Distributed ML Training Performance Engineer
Distributed ML Training Performance Engineer

OpenAI • California (MO)

Hybrid
USD 170,000 - 260,000
Relocation assistance