Virtualization & Orchestration Engineer

Jobtailor

Bellevue (WA)

On-site

USD 150,000 - 210,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Jobtailor in Bellevue, WA is seeking a platform engineer to design and operate GPU-focused infrastructure, Kubernetes orchestration, and multi-tenant provisioning for scalable AI and HPC workloads.

You will collaborate across hardware, networking, and AI platform teams to align orchestration with real-world cluster architectures, while driving reliability, security, and operational maturity of the infrastructure.

Qualifications

  • Hands-on experience with Kubernetes in production
  • Design and operate large-scale infrastructure platforms
  • Background with virtualization for cloud/HPC/GPU
  • Understanding GPU provisioning and workload scheduling

Responsibilities

  • Design and build virtualization infrastructure supporting GPU-intensive AI and HPC workloads.
  • Develop and operate Kubernetes-based orchestration systems for GPU cluster provisioning and workload scheduling.
  • Build automated provisioning systems that enable GPU capacity to be allocated, scaled, and reclaimed efficiently across multiple tenants.
  • Design solutions for workload placement, resource management, and cluster lifecycle operations.
  • Partner with hardware, networking, infrastructure, and AI platform teams to ensure orchestration systems align with real-world cluster architectures and constraints.
  • Improve the reliability, security, scalability, and operational maturity of the orchestration platform.
  • Build tooling and automation that simplifies infrastructure management and improves developer and customer experiences.
  • Contribute to architectural decisions, engineering standards, and best practices as the platform evolves.

Skills

Kubernetes orchestration
Linux systems
Infrastructure as Code
Distributed systems
Resource management

Tools

NVIDIA GPU Operator
Slurm
Container orchestration
Terraform
Cloud providers

Job description

  • Design and build virtualization infrastructure supporting GPU-intensive AI and HPC workloads.
  • Develop and operate Kubernetes-based orchestration systems for GPU cluster provisioning and workload scheduling.
  • Build automated provisioning systems that enable GPU capacity to be allocated, scaled, and reclaimed efficiently across multiple tenants.
  • Design solutions for workload placement, resource management, and cluster lifecycle operations.
  • Partner closely with hardware, networking, infrastructure, and AI platform teams to ensure orchestration systems align with real-world cluster architectures and constraints.
  • Improve the reliability, security, scalability, and operational maturity of the orchestration platform.
  • Build tooling and automation that simplifies infrastructure management and improves developer and customer experiences.
  • Contribute to architectural decisions, engineering standards, and best practices as the platform evolves.
Requirements
  • Strong hands-on experience with Kubernetes and container orchestration in production environments.
  • Experience designing, building, and operating large-scale infrastructure platforms.
  • Background with virtualization technologies supporting cloud, HPC, GPU, or distributed computing environments.
  • Understanding of GPU cluster provisioning, workload scheduling, and resource management.
  • Experience with Linux-based infrastructure and distributed systems concepts.
  • Ability to independently own complex systems from design through production operation.
  • Comfortable working in a fast-moving environment where architecture and processes are being established.
  • Experience with GPU scheduling technologies such as Slurm, Kubernetes device plugins, NVIDIA GPU Operator, or similar frameworks is preferred.
  • Experience supporting AI infrastructure, machine learning platforms, HPC environments, or GPU cloud providers is preferred.
  • Background building multi-tenant infrastructure platforms for cloud providers or large-scale compute environments is preferred.
  • Experience with infrastructure automation, Infrastructure as Code, and platform engineering practices is preferred.
  • Familiarity with high-performance networking and GPU cluster architectures is preferred.
Core Competencies

Demonstrates expertise in designing and operating Kubernetes-based orchestration systems for GPU-intensive AI and HPC workloads, with a strong focus on infrastructure automation and resource management. Proven ability to enhance the reliability and scalability of orchestration platforms while collaborating with cross-functional teams.

Highest-signal resume keywords
  • Kubernetes Orchestration
  • GPU Cluster Provisioning
  • Infrastructure Automation
  • Linux-Based Infrastructure
  • Distributed Systems
ATS Optimization Keywords
Hard Skills
  • Kubernetes
  • Virtualization Technologies
  • GPU Scheduling Technologies
  • Infrastructure as Code
  • Resource Management
  • Workload Scheduling
  • Cluster Lifecycle Operations
  • AI Infrastructure Support
  • Multi-Tenant Infrastructure
  • High-Performance Networking
Soft Skills
  • Problem-Solving
  • Collaboration
  • Adaptability
Industry Keywords
  • AI Workloads
  • HPC Workloads
  • Distributed Computing
  • Infrastructure Platforms
  • Operational Maturity
Tools & Technologies
  • Slurm
  • NVIDIA GPU Operator
  • Container Orchestration
  • Cloud Providers
  • HPC Environments
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Member of Technical Staff – AI Cloud Infrastructure
Member of Technical Staff – AI Cloud Infrastructure

Jobtailor • California (MO)

On-site
USD 150,000 - 190,000
Staff Engineer, Senior Manager
Staff Engineer, Senior Manager

Jobtailor • Connecticut

On-site
USD 140,000 - 190,000
Senior Systems Engineer, Virtualization
Senior Systems Engineer, Virtualization

Jobtailor • California (MO)

On-site
USD 150,000 - 210,000
Senior Linux Systems Engineer: Virtualization for AI
Senior Linux Systems Engineer: Virtualization for AI

Jobtailor • California (MO)

On-site
USD 150,000 - 210,000
Cluster Engineer
Cluster Engineer

STN Inc • San Francisco (CA)

On-site
USD 180,000 - 240,000
GPUaaS Kubernetes Platform Engineer
GPUaaS Kubernetes Platform Engineer

Veriipro • Irving (TX)

On-site
USD 140,000 - 180,000
Virtualization & Orchestration Engineer
Virtualization & Orchestration Engineer

Designworks Talent • Bellevue (WA)

Hybrid
USD 140,000 - 210,000
Competitive pay
Bonuses & incentives
Benefits: medical/dental/vision/401k
Senior Kubernetes Engineer
Senior Kubernetes Engineer

GTN Technical Staffing • Dallas (TX)

Hybrid
USD 150,000 - 210,000
Performance bonus
Platform Engineer
Platform Engineer

Harrison Clarke • San Francisco (CA)

On-site
USD 120,000 - 160,000
HPC Solution Architect
HPC Solution Architect

Coda Search│Staffing • Dallas (TX)

Hybrid
USD 120,000 - 160,000