Senior AI Training Performance Engineer (GPU & Scale)

figure.ai

San Jose, Northern (CA, KY)

Hybrid

USD 200,000 - 400,000

Full time

10 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Figure is seeking an experienced AI Training Performance Engineer to advance distributed training for massive models. The role focuses on optimizing GPU kernels, accelerator selection, and co-designing models to maximize hardware utilization.

You will develop kernels, build performance dashboards, and collaborate across teams to enable scalable, reliable training at scale. Applicants should have 3+ years in AI performance engineering, deep GPU knowledge, and strong Python and CUDA/C++ skills,

Qualifications

  • 3+ years in AI performance engineering with large-scale projects.
  • Deep understanding of GPU architecture and performance metrics.
  • Strong Python and CUDA/C++ skills; capable of reading framework internals.
  • Experience with profiling and translating traces into optimizations.
  • Familiarity with high-performance interconnects (NCCL, RDMA, NVLink).

Responsibilities

  • Optimize training performance for 100B+ parameter models across 100k+ GPUs.
  • Collaborate on accelerator choice, cluster topology, scheduling, and hardware procurement.
  • Write and optimize custom kernels (Triton/CUDA).
  • Build tooling and dashboards for performance monitoring and regression detection.
  • Improve data loading, checkpointing, fault tolerance, and elastic restart.
  • Co-design model architectures and training recipes for scale.

Skills

GPU architecture
Profiling tools
Python
CUDA/C++
NCCL
Distributed training
Performance engineering

Education

Bachelor's or Master's degree in CS/EE

Tools

Nsight Systems/Compute
PyTorch Profiler
HTA

Job description

Figure is seeking an experienced AI Training Performance Engineer to advance distributed training for massive models. The role focuses on optimizing GPU kernels, accelerator selection, and co-designing models to maximize hardware utilization.

You will develop kernels, build performance dashboards, and collaborate across teams to enable scalable, reliable training at scale. Applicants should have 3+ years in AI performance engineering, deep GPU knowledge, and strong Python and CUDA/C++ skills,

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI Training Infra Engineer - Scale GPU Clusters
Senior AI Training Infra Engineer - Scale GPU Clusters

Designworks Talent • Bellevue (WA)

Hybrid
USD 180,000 - 240,000
Medical insurance
401(k) with company match
Paid holidays
Senior AI Inference Performance Engineer — Scale GPUs
Senior AI Inference Performance Engineer — Scale GPUs

NVIDIA • California (MO)

On-site
USD 124,000 - 196,000
Equity eligibility
GPU Performance Engineer: Scale ML Inference & Systems
GPU Performance Engineer: Scale ML Inference & Systems

Anthropic • New York (NY)

Hybrid
USD 280,000 - 850,000
Research Engineer, GPU Performance
Research Engineer, GPU Performance

Harnham • California (MO)

On-site
USD 120,000 - 160,000
Remote GPU Performance Engineer: Scale Training & Inference
Remote GPU Performance Engineer: Scale Training & Inference

Reka • United States

Remote
USD 120,000 - 150,000
Five weeks of paid leave
Comprehensive healthcare benefits
Visa support for H1B and OPT transfers
Senior Training Infra Engineer - 800+ GPU Scale
Senior Training Infra Engineer - 800+ GPU Scale

Figureai • San Jose (CA)

On-site
USD 150,000 - 350,000
Senior AI Performance Engineer – GPU & DL
Senior AI Performance Engineer – GPU & DL

Jobtailor • California (MO)

On-site
USD 180,000 - 280,000
Senior AI Infrastructure Engineer — Scale GPU Clusters Remote
Senior AI Infrastructure Engineer — Scale GPU Clusters Remote

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 280,000 - 420,000
Equity options
Health, vision, dental benefits
Unlimited PTO
+2
Senior AI Performance & Efficiency Engineer - Equity Eligible
Senior AI Performance & Efficiency Engineer - Equity Eligible

NVIDIA • California (MO)

On-site
USD 152,000 - 288,000
Equity
Competitive benefits
Helix AI Engineer, Training Performance
Helix AI Engineer, Training Performance

figure.ai • San Jose (CA), Northern (KY)

Hybrid
USD 200,000 - 400,000