Senior AI Training Infra Engineer - Scale GPU Clusters

Designworks Talent

Bellevue (WA)

Hybrid

USD 180,000 - 240,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Medical insurance
401(k) with company match
Paid holidays

Job summary

Designworks Talent is hiring an AI Training Infrastructure Engineer to build and scale the distributed systems powering large-scale AI model training. The role focuses on reliability, efficiency, and operational excellence across GPU clusters to enable researchers and engineers to train and deploy advanced models at scale.

You will collaborate with platform and ML teams, own technically challenging projects, and contribute to the evolution of the AI infrastructure platform in a fast-moving

Qualifications

  • Experience with distributed training systems and multi-node GPU infrastructures.
  • Experience integrating training systems with production ML pipelines.
  • Strong programming skills and ownership of complex projects.

Responsibilities

  • Build and scale distributed training infrastructure for large AI models across GPU clusters.
  • Design systems to improve training reliability, efficiency, and resource utilization.
  • Develop fault tolerance, checkpointing, recovery, and large-scale training workflows.
  • Integrate AI models into production training pipelines with platform teams.
  • Diagnose and resolve training throughput, stability, and cost issues.
  • Create tooling and automation to improve researcher/developer experience.
  • Establish best practices for training infra and platform reliability.
  • Contribute to evolution of AI infra platform as an early team member.

Skills

Distributed systems
GPU training
Python
Ownership

Tools

Kubernetes
PyTorch Distributed
DeepSpeed
Megatron-LM
Ray

Job description

Designworks Talent is hiring an AI Training Infrastructure Engineer to build and scale the distributed systems powering large-scale AI model training. The role focuses on reliability, efficiency, and operational excellence across GPU clusters to enable researchers and engineers to train and deploy advanced models at scale.

You will collaborate with platform and ML teams, own technically challenging projects, and contribute to the evolution of the AI infrastructure platform in a fast-moving

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI Infrastructure Engineer — Scale GPU Clusters Remote
Senior AI Infrastructure Engineer — Scale GPU Clusters Remote

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 280,000 - 420,000
Equity options
Health, vision, dental benefits
Unlimited PTO
+2
Senior AI Infrastructure Engineer — Scale GPU Clusters
Senior AI Infrastructure Engineer — Scale GPU Clusters

AI Breaking Wire • San Francisco (CA)

On-site
USD 280,000 - 400,000
Equity
Medical, dental, and vision benefits
Unlimited PTO
+2
Senior AI Infrastructure Engineer | Scale GPU Clusters
Senior AI Infrastructure Engineer | Scale GPU Clusters

Fuel Talent LLC • Seattle (WA)

Hybrid
USD 126,000 - 189,000
Senior ML Infra Engineer - Scale GPU Clusters, Remote
Senior ML Infra Engineer - Scale GPU Clusters, Remote

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 320,000 - 500,000
Equity
Medical/Dental/Vision coverage
Unlimited PTO
+1
Senior AI Training Performance Engineer (GPU & Scale)
Senior AI Training Performance Engineer (GPU & Scale)

figure.ai • San Jose (CA), Northern (KY)

Hybrid
USD 200,000 - 400,000
Senior AI Infrastructure Engineer — GPU Clusters
Senior AI Infrastructure Engineer — GPU Clusters

Nvidia Corporation • Santa Clara (CA)

On-site
USD 152,000 - 288,000
Equity
Benefits
Senior Training Infra Engineer - 800+ GPU Scale
Senior Training Infra Engineer - 800+ GPU Scale

Figureai • San Jose (CA)

On-site
USD 150,000 - 350,000
Senior Remote ML Infrastructure Engineer: GPU & Scale
Senior Remote ML Infrastructure Engineer: GPU & Scale

Bright Vision Technologies • Bellevue (WA)

On-site
USD 100,000 - 150,000
Senior DL Infra Engineer - Multi-GPU AI Training
Senior DL Infra Engineer - Multi-GPU AI Training

2100 NVIDIA USA • California (MO)

On-site
USD 272,000 - 431,000
Staff AI Infra Engineer: Scale GPU AI Platforms
Staff AI Infra Engineer: Scale GPU AI Platforms

Seekr • San Francisco (CA)

Hybrid
USD 180,000 - 260,000
Equity Ownership – RSUs
Unlimited PTO + 14 paid holidays
Flexible hybrid work environment
+2