Helix AI Engineer, Training Performance

figure.ai

San Jose, Northern (CA, KY)

On-site

USD 200,000 - 400,000

Full time

7 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Figure is seeking an experienced AI Training Performance Engineer to advance distributed training for massive models. The role focuses on optimizing GPU kernels, accelerator selection, and co-designing models to maximize hardware utilization.

You will develop kernels, build performance dashboards, and collaborate across teams to enable scalable, reliable training at scale. Applicants should have 3+ years in AI performance engineering, deep GPU knowledge, and strong Python and CUDA/C++ skills,

Qualifications

  • 3+ years in AI performance engineering with large-scale projects.
  • Deep understanding of GPU architecture and performance metrics.
  • Strong Python and CUDA/C++ skills; capable of reading framework internals.
  • Experience with profiling and translating traces into optimizations.
  • Familiarity with high-performance interconnects (NCCL, RDMA, NVLink).

Responsibilities

  • Optimize training performance for 100B+ parameter models across 100k+ GPUs.
  • Collaborate on accelerator choice, cluster topology, scheduling, and hardware procurement.
  • Write and optimize custom kernels (Triton/CUDA).
  • Build tooling and dashboards for performance monitoring and regression detection.
  • Improve data loading, checkpointing, fault tolerance, and elastic restart.
  • Co-design model architectures and training recipes for scale.

Skills

GPU architecture
Profiling tools
Python
CUDA/C++
NCCL
Distributed training
Performance engineering

Education

Bachelor's or Master's degree in CS/EE

Tools

Nsight Systems/Compute
PyTorch Profiler
HTA

Job description

Figure is an AI robotics company developing autonomous general-purpose humanoid robots. The goal of the company is to ship humanoid robots with human level intelligence. Its robots are engineered to perform a variety of tasks in the home and commercial markets. Figure is headquartered in San Jose, CA.

Figure's vision is to deploy autonomous humanoids at a global scale. Our Helix team is looking for an experienced AI Training Performance Engineer to take our model training to the next level. This role is focused on improving distributed training frameworks for large scale model training, optimizing GPU kernels, exploring the relative gains of different accelerator types and co-designing our models to maximize utilization of our hardware.

Responsibilities

  • Optimize training performance for a 100B+ parameter models across 100k+ GPUs.
  • Collaborate with the broader team on accelerator choice, cluster topology, scheduling, and hardware procurement decisions to inform future scaling.
  • Write and optimize custom kernels (Triton/CUDA)
  • Build tooling and dashboards for continuous performance monitoring, regression detection, and root-cause analysis across training jobs
  • Optimize data loading and preprocessing pipelines so I/O never gates the accelerators
  • Improve checkpointing, fault tolerance, and elastic restart so large jobs recover quickly from node failures without losing significant wall-clock time
  • Partner with researchers to co-design model architectures and training recipes that are performant at scale (e.g., activation checkpointing strategies, mixed precision, sequence packing)
  • Extend and contribute to kernel compilers (e.g., Triton, Gluon) to improve iteration speed and enable targeting of custom/non-NVIDIA accelerators
  • Build and extend agentic systems that automatically generate, benchmark, and iterate on custom kernels
  • Evaluate emerging accelerator architectures (AMD, TPU, SRAM-based ASICs, and other novel hardware) for fit with our training workloads, and lead proof-of-concept ports/benchmarks

Requirements

  • Bachelor's or Master's degree in Computer Science, Computer/Electrical Engineering, or a related field
  • 3+ years in AI performance engineering, with significant time leading large-scale performance improvement projects
  • Deep understanding of GPU architecture and performance characteristics (memory bandwidth, compute-bound vs. memory-bound ops, occupancy)
  • Proficiency with profiling tools (Nsight Systems/Compute, PyTorch Profiler, HTA, or similar) and ability to translate traces into concrete optimizations
  • Solid grasp of collective communication (NCCL) and modern networking concepts (RDMA, NVLink, InfiniBand/RoCE, topology-aware placement).
  • Strong Python and CUDA/C++ skills; comfortable reading and modifying framework internals
  • Experience debugging performance regressions and instability at scale (stragglers, hangs, OOMs, numerical divergence)
  • Experience defining and reasoning about hardware-efficiency metrics (MFU/HFU) and using them to drive optimization priorities

Bonus Qualifications

  • Experience with heterogeneous or multi-datacenter training setups and cross-cluster orchestration
  • Contributions to open-source ML systems projects (PyTorch, Megatron-LM, vLLM, DeepSpeed, JAX, etc.)
  • Exposure to non-NVIDIA accelerators (AMD GPUs, TPU/Trainium/Inferentia, or custom silicon) and heterogeneous fleet management.

The US base salary range for this full-time position is between $200,000 - $400,000 annually.

The pay offered for this position may vary based on several individual factors, including job-related knowledge, skills, and experience. The total compensation package may also include additional components/benefits depending on the specific role. This information will be shared if an employment offer is extended.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Helix AI Engineer, Training Infrastructure
Helix AI Engineer, Training Infrastructure

Figureai • San Jose (CA)

On-site
USD 150,000 - 350,000
Staff AI Inference and Acceleration Engineer
Staff AI Inference and Acceleration Engineer

Figureai • San Jose (CA)

On-site
USD 180,000 - 275,000
Helix AI Engineer, Pretraining
Helix AI Engineer, Pretraining

Figureai • San Jose (CA)

On-site
USD 120,000 - 160,000
Staff AI Inference and Acceleration Engineer
Staff AI Inference and Acceleration Engineer

Figure • San Jose (CA)

On-site
USD 180,000 - 275,000
Helix AI Engineer, Modeling
Helix AI Engineer, Modeling

Figureai • San Jose (CA)

On-site
USD 120,000 - 150,000
Helix AI Engineer, Pretraining
Helix AI Engineer, Pretraining

Figure • San Jose (CA)

On-site
USD 120,000 - 150,000
Machine Learning Performance Engineer – Offboard Training & Inference
Machine Learning Performance Engineer – Offboard Training & Inference

NLP PEOPLE • Sunnyvale (CA)

On-site
USD 215,000 - 285,000
Helix AI Engineer, Video Pretraining
Helix AI Engineer, Video Pretraining

Figureai • San Jose (CA)

On-site
USD 120,000 - 160,000
Helix AI Engineer, Data Infrastructure
Helix AI Engineer, Data Infrastructure

Figureai • San Jose (CA)

On-site
USD 150,000 - 350,000
Helix AI Engineer, Data Infrastructure
Helix AI Engineer, Data Infrastructure

Figure • San Jose (CA)

On-site
USD 150,000 - 350,000