Senior AI Training Infra Engineer — GPU Clusters

Designworks Talent

Bellevue (WA)

Hybrid

USD 180,000 - 300,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Medical, dental, vision insurance
401(k) with company match
Paid holidays

Job summary

Designworks Talent is seeking an AI Training Infrastructure Engineer for a hybrid role in the Bellevue, WA area. You will build and scale distributed training infrastructure powering large AI models across multi-node GPU clusters and production pipelines.

Work with platform, orchestration, performance, and ML teams to improve reliability, efficiency, and scalability. Responsibilities include fault-tolerance, checkpointing, recovery, and developing developer tools to streamline AI research

Qualifications

  • Hands-on experience building distributed training systems.
  • Experience with large AI models and production ML pipelines.
  • Ability to optimize multi-node GPU training for reliability and efficiency.
  • Proven track record delivering scalable ML infrastructure solutions.

Responsibilities

  • Build and scale distributed training infrastructure supporting large AI models across GPU clusters.
  • Design and improve systems for training reliability, efficiency, and resource utilization.
  • Develop fault tolerance, checkpointing, and recovery for large-scale training.
  • Integrate training systems with production ML pipelines in cross-functional teams.
  • Diagnose and resolve bottlenecks affecting throughput, stability, and cost efficiency.
  • Create tools and automation to improve researcher and engineer productivity.
  • Establish best practices for training infra and platform reliability.
  • Contribute to evolving the AI infra platform as an early engineering team member.

Skills

Distributed training
GPU clusters
Python
Multi-node systems
Production ML pipelines

Tools

PyTorch Distributed
DeepSpeed
Kubernetes
GPU virtualization

Job description

Designworks Talent is seeking an AI Training Infrastructure Engineer for a hybrid role in the Bellevue, WA area. You will build and scale distributed training infrastructure powering large AI models across multi-node GPU clusters and production pipelines.

Work with platform, orchestration, performance, and ML teams to improve reliability, efficiency, and scalability. Responsibilities include fault-tolerance, checkpointing, recovery, and developing developer tools to streamline AI research

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI Infrastructure Engineer | Scale GPU Clusters
Senior AI Infrastructure Engineer | Scale GPU Clusters

Fuel Talent LLC • Seattle (WA)

Hybrid
USD 126,000 - 189,000
Senior AI Infrastructure Engineer — Scale GPU Clusters Remote
Senior AI Infrastructure Engineer — Scale GPU Clusters Remote

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 280,000 - 420,000
Equity options
Health, vision, dental benefits
Unlimited PTO
+2
Senior Training Infra Engineer - 800+ GPU Scale
Senior Training Infra Engineer - 800+ GPU Scale

Figureai • San Jose (CA)

On-site
USD 150,000 - 350,000
Senior AI Infrastructure Engineer — Scale GPU Clusters
Senior AI Infrastructure Engineer — Scale GPU Clusters

AI Breaking Wire • San Francisco (CA)

On-site
USD 280,000 - 400,000
Equity
Medical, dental, and vision benefits
Unlimited PTO
+2
Staff AI Training Infrastructure Engineer
Staff AI Training Infrastructure Engineer

Designworks Talent • Bellevue (WA)

Hybrid
USD 180,000 - 300,000
Medical, dental, vision insurance
401(k) with company match
Paid holidays
Senior AI Infrastructure Lead - GPU Clusters & Model Serving
Senior AI Infrastructure Lead - GPU Clusters & Model Serving

Outsourceit • San Francisco (CA)

On-site
USD 120,000 - 170,000
Senior AI Infra Engineer-Distributed GPU Clusters (Equity)
Senior AI Infra Engineer-Distributed GPU Clusters (Equity)

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 184,000 - 356,500
Senior AI Infra Architect for Multi-GPU Training + Equity
Senior AI Infra Architect for Multi-GPU Training + Equity

NVIDIA Gruppe • California (MO)

On-site
USD 224,000 - 431,000
Equity
Benefits
AI/ML Infra Engineer — GPU Clusters & HPC
AI/ML Infra Engineer — GPU Clusters & HPC

NVIDIA AI • Redmond (WA)

On-site
USD 120,000 - 180,000
Competitive salaries
Comprehensive benefits package
Equity
Senior AI Infra Engineer: GPU Clusters & Kubernetes
Senior AI Infra Engineer: GPU Clusters & Kubernetes

Intelliswift - An LTTS Company • Sunnyvale (CA)

On-site
USD 120,000 - 150,000