Staff AI Training Infrastructure Engineer

Designworks Talent LLC

Bellevue (KY)

Hybrid

USD 150,000 - 210,000

Full time

5 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Medical, dental, vision insurance
401(k) with company match
Paid holidays

Job summary

Designworks Talent LLC in Bellevue, WA is seeking an AI Training Infrastructure Engineer to build and scale distributed training systems for large AI models, operating across multi-node GPU clusters.

You will collaborate with platform, orchestration, and performance teams to improve reliability, efficiency, and training throughput while enabling researchers to push the boundaries of AI workloads.

Qualifications

  • Hands-on experience building and operating distributed training systems or large-scale ML infrastructure.
  • Experience supporting large AI models, foundation models, post-training workflows, or similar ML systems.
  • Strong understanding of reliability, scalability, and efficiency challenges with multi-node GPU training.
  • Experience integrating training systems with production ML pipelines.

Responsibilities

  • Build and scale distributed training infrastructure for large AI models across GPU clusters.
  • Design and improve systems for reliability, efficiency, and resource utilization.
  • Develop fault tolerance, checkpointing, recovery for large-scale training operations.
  • Integrate AI models into production training pipelines with platform and performance teams.
  • Diagnose and resolve issues impacting training throughput, stability, and cost efficiency.
  • Build tools and automation to improve developer experience for AI researchers and engineers.
  • Establish best practices for training infrastructure and platform reliability.
  • Contribute to evolution of AI infrastructure platform as an early engineering team member.

Skills

Distributed training
ML infrastructure
Reliability engineering
Production ML pipelines
Distributed systems
Programming skills

Tools

Kubernetes
Containerization

Job description

Staff AI Training Infrastructure Engineer

Location: Hybrid | Bellevue, WA Areamultiple roles available

Build the Training Infrastructure Powering Next-Generation AI Models

About the Opportunity

A well-funded, rapidly growing AI infrastructure company is building a next-generation cloud platform designed to power the full lifecycle of artificial intelligence. The organization is developing a comprehensive AI infrastructure, platform, and services portfolio that supports the full spectrum of AI workloads---including large-scale compute, model training, fine-tuning, inference, and emerging agentic AI applications.

Backed by significant long-term investment, the company combines the speed, ownership, and innovation of a startup with the stability and resources of an established parent organization. Engineering teams are intentionally lean, highly collaborative, and AI-native, leveraging modern tooling and automation to build infrastructure capable of supporting the industry's most demanding AI workloads.

We're seeking AI Training Infrastructure Engineers to build and scale the distributed systems that power large-scale AI model training. This team focuses on reliability, efficiency, and operational excellence across GPU clusters, enabling researchers and engineers to train and deploy advanced AI models at scale.

The Opportunity

This is a foundational engineering role focused on building the infrastructure layer behind large-scale AI training workloads. You'll work on distributed training systems, GPU clusters, model pipelines, and the tooling required to make AI development more reliable, efficient, and scalable.

You will collaborate closely with infrastructure, orchestration, performance, and machine learning teams to solve complex challenges around distributed computing, fault tolerance, training efficiency, and production readiness.

This opportunity is ideal for engineers who enjoy building highly scalable systems and working at the intersection of AI research, infrastructure engineering, and distributed computing.

What You'll Do
  • Build and scale distributed training infrastructure supporting large AI models across large GPU clusters.

  • Design and improve systems that increase training reliability, efficiency, and resource utilization.

  • Develop solutions for fault tolerance, checkpointing, recovery, and large-scale training operations.

  • Integrate AI models into production training pipelines in partnership with platform, orchestration, and performance engineering teams.

  • Diagnose and resolve issues impacting training throughput, stability, reliability, and cost efficiency.

  • Build tools and automation that improve the developer experience for AI researchers and engineers.

  • Establish best practices for training infrastructure, operational processes, and platform reliability.

  • Contribute to the evolution of the AI infrastructure platform as an early member of the engineering team.

What We're Looking For
  • Hands‑on experience building and operating distributed training systems or large‑scale machine learning infrastructure.

  • Experience supporting large AI models, foundation models, post‑training workflows, or similar ML systems.

  • Strong understanding of the reliability, scalability, and efficiency challenges associated with multi‑node GPU training.

  • Experience integrating training systems with production machine learning pipelines.

  • Strong programming skills and experience working with complex distributed systems.

  • Ability to independently own technically challenging projects in a fast‑moving engineering environment.

  • Comfortable operating with high ownership and limited process overhead.

Preferred Qualifications
  • Experience with distributed training frameworks such as PyTorch Distributed, DeepSpeed, Megatron‑LM, Ray, or similar technologies.

  • Experience with supervised fine‑tuning (SFT), reinforcement learning from human feedback (RLHF), or other post‑training workflows.

  • Background operating AI training infrastructure at scale within a hyperscaler, AI research organization, cloud provider, or GPU cloud environment.

  • Experience optimizing GPU utilization, training performance, or distributed system reliability.

  • Familiarity with Kubernetes, containerized AI workloads, and large‑scale infrastructure platforms.

Compensation
  • Competitive base pay for Bellevue market

  • Certain roles are eligible for additional rewards, including merit increases, annual bonus, and long term incentives. These awards are allocated based on individual performance

  • U.S. based employees have access to medical, dental, and vision insurance, a 401(k) plan and company match, employees also receive per calendar year, paid holidays

Location
  • Hybrid role based in the Bellevue, WA area.

  • Approximately three days per week in the office.

  • Candidates elsewhere in the U.S. who are open to relocation are encouraged to apply.

  • U.S. work authorization is required. Visa sponsorship is not currently available.

Why Join?
  • Build the infrastructure powering the next generation of AI models and applications.

  • Work directly on distributed training systems, GPU clusters, and large-scale AI platforms.

  • Solve some of the industry's most challenging problems around AI scalability, reliability, and efficiency.

  • Join early enough to influence architecture, tooling, and engineering practices.

  • Collaborate with a highly experienced team building critical AI infrastructure from the ground up.

  • Enjoy the ownership and technical impact of a startup environment backed by significant long-term investment.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior AI Training Infrastructure Engineer
Senior AI Training Infrastructure Engineer

Designworks Talent LLC • Bellevue (WA)

Hybrid
USD 180,000 - 260,000
Medical Insurance
Dental Insurance
Vision Insurance
+2
Principal Data Center Operations and Maintenance Engineer
Principal Data Center Operations and Maintenance Engineer

Designworks Talent LLC • Bellevue (WA)

Hybrid
USD 150,000 - 230,000
Medical, dental, and vision insurance
401(k) plan with company match
Paid holidays
Staff Principal Data Center Operations and Maintenance Engineer
Staff Principal Data Center Operations and Maintenance Engineer

Designworks Talent LLC • Bellevue (KY)

Hybrid
USD 120,000 - 170,000
Medical Insurance
401(k) Match
Paid Holidays
+1
Staff Principal Data Center Operations and Maintenance Engineer
Staff Principal Data Center Operations and Maintenance Engineer

Designworks Talent LLC • Bellevue (WA)

Hybrid
USD 130,000 - 200,000
Medical, dental, vision insurance
401(k) with company match
Paid holidays
Staff Inference Engineer
Staff Inference Engineer

Designworks Talent LLC • Bellevue (KY)

Hybrid
USD 180,000 - 240,000
Health insurance
401(k) plan with company match
Paid holidays
Staff Senior Virtualization & Orchestration Engineer
Staff Senior Virtualization & Orchestration Engineer

Designworks Talent LLC • Bellevue (WA)

Hybrid
USD 150,000 - 210,000
Medical, dental, vision insurance
401(k) plan with company match
Paid holidays
Member of Technical Staff - Training Platform
Member of Technical Staff - Training Platform

Prime Intellect AI • San Francisco (CA)

Hybrid
USD 150,000 - 300,000
Remote or SF office
Visa sponsorship
Relocation support
+3
Principal Data Center Operations and Maintenance Engineer
Principal Data Center Operations and Maintenance Engineer

Designworks Talent LLC • Bellevue (KY)

Hybrid
USD 120,000 - 180,000
Medical, dental, and vision insurance
401(k) plan with company match
Paid holidays
Data Center Operations and Maintenance Engineering Leader
Data Center Operations and Maintenance Engineering Leader

Designworks Talent • Bellevue (WA)

On-site
USD 180,000 - 240,000
Health insurance
Vision coverage
401(k) plan
Staff Data Center Infrastructure Software Engineer
Staff Data Center Infrastructure Software Engineer

Designworks Talent LLC • Bellevue (WA)

Hybrid
USD 140,000 - 200,000
Merit-based bonus
Long-term incentives
Medical, dental, vision