Helix AI Engineer, Training Infrastructure

Figureai

San Jose (CA)

On-site

USD 150,000 - 350,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Figureai in San Jose, CA is seeking an experienced Training Infrastructure Engineer to manage and optimize their training infrastructure for AI robotics. You will collaborate with AI researchers to scale the training of new model architectures and enhance performance through distributed training algorithms.

A strong background in Python and deep learning frameworks is essential, along with experience managing large-scale GPU clusters. Join an innovative team and help lead the future of humanoid robotics.

Qualifications

  • Strong software engineering fundamentals.
  • Extensive professional experience with Python and PyTorch.
  • Proven track record of scaling and running large-scale training experiments on 800+ GPUs.
  • Minimum of 4 years of professional experience in building reliable infrastructure.

Responsibilities

  • Design, deploy, and maintain training clusters.
  • Optimize and maintain scalable deep learning frameworks.
  • Implement distributed training and data loaders.
  • Profile, identify, and eliminate training bottlenecks.
  • Implement tooling for data processing and model experimentation.

Skills

Python
PyTorch
Backend systems
GPU management
HPC clusters

Education

Bachelor's or Master's degree in Computer Science, Robotics, Engineering, or related field

Tools

AWS
Azure
GCP
CUDA
SLURM
Kubernetes

Job description

Figure is an AI robotics company developing autonomous general-purpose humanoid robots. The goal of the company is to ship humanoid robots with human level intelligence. Its robots are engineered to perform a variety of tasks in the home and commercial markets. Figure is headquartered in San Jose, CA.

Figure's vision is to deploy autonomous humanoids at a global scale. Our Helix team is looking for an experienced Training Infrastructure Engineer to take our infrastructure to the next level. This role is focused on managing the training cluster, implementing distributed training algorithms, data loaders, and developer tools for AI researchers.

Responsibilities
  • Design, deploy, and maintain Figure's training clusters
  • Architect, optimize, and maintain scalable deep learning frameworks for training on massive robot datasets
  • Work together with AI researchers to implement training of new model architectures at a large scale
  • Implement distributed training, advanced parallelization strategies, and high-performance data loaders to reduce model development cycles
  • Profile, identify, and eliminate training bottlenecks at the hardware and software levels to maximize Model FLOPs Utilization (MFU)
  • Implement tooling for data processing, model experimentation, and continuous integration
Requirements
  • Strong software engineering fundamentals
  • Bachelor's or Master's degree in Computer Science, Robotics, Engineering, or a related field
  • Extensive professional experience with Python and PyTorch
  • Proven track record of scaling and running large-scale training experiments personally on 800+ GPUs
  • Experience managing HPC clusters for deep neural network training
  • Minimum of 4 years of professional, full-time experience building reliable backend systems and infrastructure
Bonus Qualifications
  • Experience contributing to or maintaining open-source distributed training frameworks (Megatron-LM, DeepSpeed, TorchTitan)
  • Experience managing cloud infrastructure (AWS, Azure, GCP)
  • Experience with job scheduling / orchestration tools (SLURM, Kubernetes, LSF, etc.)
  • Experience with configuration management tools (Ansible, Terraform, Puppet, Chef, etc.)
  • Deep understanding of CUDA and hands‑on experience writing custom GPU kernels to optimize training

The US base salary range for this full-time position is between $150,000 - $350,000 annually.

The pay offered for this position may vary based on several individual factors, including job-related knowledge, skills, and experience. The total compensation package may also include additional components/benefits depending on the specific role. This information will be shared if an employment offer is extended.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Helix AI Engineer, Training Performance
Helix AI Engineer, Training Performance

figure.ai • San Jose (CA), Northern (KY)

Hybrid
USD 200,000 - 400,000
Helix AI Engineer, Data Infrastructure
Helix AI Engineer, Data Infrastructure

Figure • San Jose (CA)

On-site
USD 150,000 - 350,000
Helix AI Engineer, Data Infrastructure
Helix AI Engineer, Data Infrastructure

Figureai • San Jose (CA)

On-site
USD 150,000 - 350,000
Helix AI Engineer, Pretraining
Helix AI Engineer, Pretraining

Figureai • San Jose (CA)

On-site
USD 120,000 - 160,000
Helix AI Engineer, Pretraining
Helix AI Engineer, Pretraining

Figure • San Jose (CA)

On-site
USD 120,000 - 150,000
Helix AI Engineer, Modeling
Helix AI Engineer, Modeling

Figureai • San Jose (CA)

On-site
USD 120,000 - 150,000
Helix AI Engineer, Video Pretraining
Helix AI Engineer, Video Pretraining

Figureai • San Jose (CA)

On-site
USD 120,000 - 160,000
Helix AI Engineer, Backend
Helix AI Engineer, Backend

Figureai • San Jose (CA)

On-site
USD 150,000 - 400,000
Helix AI Engineer, Perception Figure AI San Jose, CA $200,000 - $350,000/yr
Helix AI Engineer, Perception Figure AI San Jose, CA $200,000 - $350,000/yr

Neura Market • San Jose (CA), Northern (KY)

Hybrid
USD 200,000 - 350,000
Helix AI Engineer, Backend
Helix AI Engineer, Backend

Figure • San Jose (CA)

On-site
USD 150,000 - 400,000