VP – AI Infrastructure Engineering

Jobtailor

Bellevue (WA)

On-site

USD 200,000 - 350,000

Full time

5 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Jobtailor is seeking a Senior Infrastructure Leader to Bellevue, WA, driving data center infrastructure bring-up and production readiness. Lead platform lifecycle from hardware to automated provisioning and workload readiness, building scalable automation for bare-metal provisioning and Linux deployment.

You will oversee GPU clusters, Kubernetes environments, and distributed compute infrastructure, establishing standards for servers, GPUs, networking, and firmware.

Qualifications

  • 12+ years of experience across infrastructure, software, systems, platform engineering, or related technical disciplines.
  • 5+ years of engineering leadership experience, including managing and developing highly technical teams
  • Proven experience building and operating large-scale data center, cloud, HPC, or AI infrastructure
  • Strong technical understanding of Linux, distributed systems, networking, and infrastructure automation
  • Hands-on understanding of Kubernetes, containers, and infrastructure-as-code
  • Demonstrated ability to lead complex infrastructure deployments and bring new environments into production
  • Ability to operate across hardware and software organizations
  • Strong communication and cross-functional leadership skills
  • Hands-on, high-ownership leadership style suited to a fast-moving, build-from-the-ground-up environment
  • Preferred: experience with GPU infrastructure, NVIDIA platforms, AI or HPC environments

Responsibilities

  • Lead the engineering organization responsible for data center infrastructure bring-up and production readiness
  • Own the platform lifecycle from installed hardware through automated provisioning, configuration, validation, and workload readiness
  • Build and scale automation for bare-metal provisioning, Linux deployment, configuration management, and infrastructure validation
  • Lead deployment and configuration of GPU clusters, Kubernetes environments, and distributed compute infrastructure
  • Establish engineering standards for servers, GPUs, networking, storage, firmware, and system configuration
  • Drive infrastructure automation using Terraform, Ansible, Bash, and similar technologies
  • Oversee integration with Redfish, IPMI, BMCs, PXE, MAAS, Ironic, Foreman, or comparable platforms
  • Partner with Network, Hardware/GPU, Data Center Operations, SRE, and Software Engineering teams
  • Establish automated testing, health checks, monitoring, and validation processes
  • Improve deployment speed, reliability, automation, scalability, and operational efficiency
  • Build and develop a highly capable engineering organization and scalable processes

Skills

Infrastructure Automation
Kubernetes Deployment
Linux Systems Management
Engineering Leadership
Data Center Operations

Tools

Terraform
Ansible
Bash
Redfish
IPMI
BMC
PXE
MAAS
Ironic
Foreman

Job description

  • Lead the engineering organization responsible for data center infrastructure bring-up and production readiness
  • Own the platform lifecycle from installed hardware through automated provisioning, configuration, validation, and workload readiness
  • Build and scale automation for bare-metal provisioning, Linux deployment, configuration management, and infrastructure validation
  • Lead deployment and configuration of GPU clusters, Kubernetes environments, and distributed compute infrastructure
  • Establish engineering standards for servers, GPUs, networking, storage, firmware, and system configuration
  • Drive infrastructure automation using Terraform, Ansible, Bash, and similar technologies
  • Oversee integration with Redfish, IPMI, BMCs, PXE, MAAS, Ironic, Foreman, or comparable platforms
  • Partner with Network, Hardware/GPU, Data Center Operations, SRE, and Software Engineering teams
  • Establish automated testing, health checks, monitoring, and validation processes
  • Improve deployment speed, reliability, automation, scalability, and operational efficiency
  • Build and develop a highly capable engineering organization and scalable processes
Requirements
  • 12+ years of experience across infrastructure, software, systems, platform engineering, or related technical disciplines
  • 5+ years of engineering leadership experience, including managing and developing highly technical teams
  • Proven experience building and operating large-scale data center, cloud, HPC, or AI infrastructure
  • Strong technical understanding of Linux, distributed systems, networking, and infrastructure automation
  • Hands-on understanding of Kubernetes, containers, and infrastructure-as-code
  • Demonstrated ability to lead complex infrastructure deployments and bring new environments into production
  • Ability to operate across hardware and software organizations
  • Strong communication and cross-functional leadership skills
  • Hands-on, high-ownership leadership style suited to a fast-moving, build-from-the-ground-up environment
  • Preferred: experience with GPU infrastructure, NVIDIA platforms, AI or HPC environments
  • Preferred: bare-metal provisioning technologies such as MAAS, Ironic, xCAT, Foreman, or similar
  • Preferred: Redfish, IPMI, BMC, PXE, and firmware management
  • Preferred: CUDA, NVML, DCGM, NVIDIA drivers, or GPU Operator
  • Preferred: InfiniBand, RoCE, RDMA, or high-speed Ethernet
  • Preferred: Kubernetes, Slurm, or similar cluster orchestration and scheduling technologies
  • Preferred: automated infrastructure validation and hardware health testing
  • Preferred: experience scaling infrastructure across thousands of servers or GPUs
  • U.S. work authorization required
  • Visa sponsorship is not currently available
  • Willingness to relocate if currently outside the Bellevue area
Core Competencies

Demonstrates extensive experience in leading engineering teams focused on data center infrastructure, automation, and deployment processes. Proficient in managing large-scale infrastructure projects, including GPU and cloud environments, while ensuring operational efficiency and scalability.

Highest-signal resume keywords
  • Infrastructure Automation
  • Kubernetes Deployment
  • Linux Systems Management
  • Engineering Leadership
  • Data Center Operations
Hard Skills
  • Infrastructure Automation
  • Linux
  • Distributed Systems
  • Kubernetes
  • Bare-Metal Provisioning
  • GPU Infrastructure
  • Infrastructure-As-Code
  • Networking
  • Cloud Infrastructure
  • AI Infrastructure
Soft Skills
  • Strong Communication
  • Cross-Functional Leadership
  • High-Ownership Leadership
Industry Keywords
  • Data Center Infrastructure
  • HPC
  • AI
  • GPU Clusters
  • Automated Testing
  • Health Checks
  • Monitoring
  • Scalability
  • Operational Efficiency
  • Deployment Speed
Tools & Technologies
  • Terraform
  • Ansible
  • Bash
  • Redfish
  • IPMI
  • BMC
  • PXE
  • MAAS
  • Ironic
  • Foreman
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

VP - AI Infrastructure Engineering
VP - AI Infrastructure Engineering

Designworks Talent • Bellevue (WA)

Hybrid
USD 300,000 - 520,000
Senior/Staff Software Engineer, Kubernetes Infrastructure
Senior/Staff Software Engineer, Kubernetes Infrastructure

Kindredventures • United States

On-site
USD 140,000 - 190,000
Sr HPC Hardware Engineer
Sr HPC Hardware Engineer

Career Techniques • Dallas (TX)

Hybrid
USD 120,000 - 180,000
Principal Infrastructure Engineer, AI Cluster Performance & Validation
Principal Infrastructure Engineer, AI Cluster Performance & Validation

Nscale • New York (NY), San Francisco (CA), Seattle (WA)

On-site
USD 180,000 - 240,000
Cluster Engineer
Cluster Engineer

STN Inc • San Francisco (CA)

On-site
USD 180,000 - 240,000
Senior System Engineer – GPU Platforms
Senior System Engineer – GPU Platforms

Jobtailor • San Jose (CA)

On-site
USD 150,000 - 210,000
Senior Solutions Engineer, AI Infrastructure
Senior Solutions Engineer, AI Infrastructure

VAST Data • New York (NY)

On-site
USD 150,000 - 200,000
VP of Engineering
VP of Engineering

Hyperbolic • San Francisco (CA)

On-site
USD 200,000 - 300,000
AI Infrastructure Engineer
AI Infrastructure Engineer

Jobtailor • San Francisco (CA)

Hybrid
USD 140,000 - 210,000
Senior Software Engineer, Infrastructure Software for AI (Centralized AI Data Centers & Distrib[...]
Senior Software Engineer, Infrastructure Software for AI (Centralized AI Data Centers & Distrib[...]

Intelliswift - An LTTS Company • Sunnyvale (CA)

On-site
USD 120,000 - 150,000
Competitive salary
Health insurance
Flexible work hours