VP - AI Infrastructure Engineering

Designworks Talent

Bellevue (WA)

On-site

USD 300,000 - 520,000

Full time

6 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Designworks Talent seeks a Vice President of AI Infrastructure Engineering in Bellevue, WA, for a senior leadership role overseeing the engineering organization that transforms data center hardware into reliable, production-ready compute infrastructure.

The role blends deep technical expertise with organizational leadership to scale AI/HPC infrastructure across servers, GPUs, Linux, networking, and automation.

Qualifications

  • 12+ years of experience across infrastructure, software, systems, or platform engineering.
  • 5+ years of engineering leadership experience managing and developing technical teams.
  • Proven experience building and operating large-scale data center, cloud, HPC, or AI infrastructure.
  • Strong Linux, distributed systems, networking, and automation skills.

Responsibilities

  • Lead the 15+ person team for data center infrastructure bring-up and production readiness.
  • Own platform lifecycle from hardware install to automated provisioning and workload readiness.
  • Build and scale automation for bare-metal provisioning and Linux deployment.
  • Lead deployment of GPU clusters, Kubernetes environments, and distributed compute infrastructure.
  • Establish engineering standards for servers, GPUs, networking, storage, firmware.
  • Drive infrastructure automation using Terraform, Ansible, Bash.
  • Oversee integration with Redfish, IPMI, BMCs, PXE, MAAS, Ironic, Foreman.
  • Partner with Network, Hardware/GPU, DC Ops, SRE, and Software teams.
  • Establish automated testing, health checks, monitoring, validation processes.
  • Improve deployment speed, reliability, scalability and operational efficiency.
  • Build and develop a high-performing organization with scalable processes.

Skills

Infrastructure
Linux
Kubernetes
Automation
Leadership
Communication

Tools

Terraform
Ansible
Bash
MAAS

Job description

Vice President AI Infrastructure Engineering

Bellevue, WA Area | Hybrid | Senior Leadership

A rapidly growing, well-funded technology company is seeking a Vice President AI Infrastructure Engineering Leader to lead the engineering organization responsible for transforming newly deployed data center hardware into reliable, production-ready compute infrastructure.

This is a high-impact leadership opportunity for someone who combines deep technical expertise in AI/HPC infrastructure with strong engineering leadership. The ideal candidate understands both the physical infrastructure layer and the software automation required to operate large-scale GPU and compute environments efficiently.

You’ll operate at the intersection of servers, GPUs, Linux, networking, Kubernetes, distributed systems, automation, and infrastructure software, helping establish the architecture, standards, tooling, and engineering practices required to deploy and operate infrastructure at scale.

What You’ll Do
  • Lead the team of 15+ responsible for data center infrastructure bring-up and production readiness.

  • Own the platform lifecycle from installed hardware through automated provisioning, configuration, validation, and workload readiness.

  • Build and scale automation for bare-metal provisioning, Linux deployment, configuration management, and infrastructure validation.

  • Lead deployment and configuration of GPU clusters, Kubernetes environments, and distributed compute infrastructure.

  • Establish engineering standards for servers, GPUs, networking, storage, firmware, and system configuration.

  • Drive infrastructure automation using Terraform, Ansible, Bash, and similar technologies.

  • Oversee integration with technologies such as Redfish, IPMI, BMCs, PXE, MAAS, Ironic, Foreman, or comparable platforms.

  • Partner closely with Network, Hardware/GPU, Data Center Operations, SRE, and Software Engineering teams to deliver production-ready infrastructure.

  • Establish automated testing, health checks, monitoring, and validation processes to identify infrastructure issues before workloads reach production.

  • Improve deployment speed, reliability, automation, scalability, and operational efficiency.

  • Build and develop a highly capable engineering organization while establishing processes that can scale with the business.

What We’re Looking For
  • 12+ years of experience across infrastructure, software, systems, platform engineering, or related technical disciplines.

  • 5+ years of engineering leadership experience, including managing and developing highly technical teams.

  • Proven experience building and operating large-scale data center, cloud, HPC, or AI infrastructure.

  • Strong technical understanding of Linux, distributed systems, networking, and infrastructure automation.

  • Hands‑on understanding of Kubernetes, containers, and infrastructure‑as‑code.

  • Demonstrated ability to lead complex infrastructure deployments and bring new environments into production.

  • Ability to operate comfortably across both hardware and software organizations.

  • Strong communication and cross‑functional leadership skills.

  • A hands‑on, high‑ownership leadership style with the ability to operate effectively in a fast‑moving, build‑from‑the‑ground‑up environment.

Preferred Experience

Experience in one or more of the following areas is highly valued:

  • GPU infrastructure, NVIDIA platforms, AI or HPC environments.

  • Bare‑metal provisioning technologies such as MAAS, Ironic, xCAT, Foreman, or similar.

  • Hardware management technologies including Redfish, IPMI, BMC, PXE, and firmware management.

  • NVIDIA technologies such as CUDA, NVML, DCGM, NVIDIA drivers, or GPU Operator.

  • High‑performance networking including InfiniBand, RoCE, RDMA, or high‑speed Ethernet.

  • Cluster orchestration and scheduling technologies such as Kubernetes, Slurm, or similar.

  • Automated infrastructure validation and hardware health testing.

  • Experience scaling infrastructure across thousands of servers or GPUs.

The Opportunity

This is an opportunity to join an organization at an early and highly consequential stage of its growth. You’ll have significant influence over architecture, automation, engineering standards, tooling, and team development, rather than simply inheriting an established infrastructure environment.

The company operates with a startup mentality of fast, lean, highly collaborative, and high ownership while having the resources to build infrastructure for significant scale.

The role is particularly well suited to a leader who enjoys building something new, moving quickly, solving complex infrastructure challenges, and creating software‑driven systems that replace manual processes with scalable automation.

Location
  • Bellevue, WA area

  • Hybrid work model with three days per week in the office

  • Candidates currently outside the area may be considered if they are willing to relocate

  • U.S. work authorization required; visa sponsorship is not currently available

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

VP – AI Infrastructure Engineering
VP – AI Infrastructure Engineering

Jobtailor • Bellevue (WA)

On-site
USD 200,000 - 350,000
Data Center Operations and Maintenance Engineering Leader
Data Center Operations and Maintenance Engineering Leader

Designworks Talent • Bellevue (WA)

Hybrid
USD 180,000 - 240,000
Health insurance
Vision coverage
401(k) plan
Data Center Operations and Maintenance Engineer
Data Center Operations and Maintenance Engineer

Designworks Talent • Bellevue (WA)

Hybrid
USD 130,000 - 190,000
Medical insurance
Dental and vision insurance
401(k) with company match
+1
Hardware Design Engineer
Hardware Design Engineer

Designworks Talent • Bellevue (WA)

Hybrid
USD 180,000 - 240,000
Medical, dental, vision insurance
401(k) with company match
Paid holidays
VP of Engineering
VP of Engineering

Hyperbolic • San Francisco (CA)

On-site
USD 200,000 - 300,000
VP, AI Infrastructure & HPC Platform Engineering
VP, AI Infrastructure & HPC Platform Engineering

Designworks Talent • Bellevue (WA)

Hybrid
USD 300,000 - 520,000
Senior Software Engineer, Infrastructure Software for AI (Centralized AI Data Centers & Distrib[...]
Senior Software Engineer, Infrastructure Software for AI (Centralized AI Data Centers & Distrib[...]

Intelliswift - An LTTS Company • Sunnyvale (CA)

On-site
USD 120,000 - 150,000
Competitive salary
Health insurance
Flexible work hours
Head of AI Data Center Infrastructure Platforms and Software
Head of AI Data Center Infrastructure Platforms and Software

Summit Group Solutions, LLC • United States

On-site
USD 150,000 - 350,000
VP AI - Core Infrastructure
VP AI - Core Infrastructure

kadence • New York (NY)

Hybrid
USD 250,000 - 350,000
Director of Infrastructure Engineering
Director of Infrastructure Engineering

Appsierra Group • United States

On-site
USD 350,000 - 500,000
Equity compensation eligibility
Performance-based bonuses
Health insurance reimbursement up to 1
+3