Staff HPC Systems Architect

Jobtailor

California (MO)

On-site

USD 180,000 - 260,000

Full time

4 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Jobtailor is seeking a senior architect to design scalable compute platforms for AI/ML, simulation, and high-throughput workloads. You will define platform standards, map workloads to capabilities across bare metal and cloud, and guide hardware and firmware decisions.

You will lead validation and performance characterization, mentor engineers, and drive optimization of density, power, and cost. Deep expertise in CPU/GPU architectures and HPC fabrics is required.

Qualifications

  • Proven experience (7+ years) architecting large-scale 10k-100k+ GPU HPC or cloud compute platforms.
  • Deep knowledge of CPU/GPU architectures, memory hierarchies, and accelerator topologies.
  • Experience designing systems around high-bandwidth, low-latency fabrics (NVLink, InfiniBand, RoCE).
  • Strong understanding of system performance tuning, resource scheduling, thermal and power optimization, and compute lifecycle management.
  • Comfortable working across hardware and software boundaries, especially at the intersection of compute architecture, OS behavior, and orchestration layers.
  • Skilled at balancing architectural tradeoffs for density, power efficiency, cooling, and performance.
  • Strong analytical and communication skills, with a track record of influencing technical strategy across teams.
  • Strong ownership and can-do attitude, self-starter who feels comfortable working in ambiguity.
  • Hands-on experience with AI/ML workloads and their compute performance characteristics (nice to have).
  • Familiarity with orchestration tools used in HPC, such as Slurm and Kubernetes (nice to have).
  • Experience with virtualization technologies, specifically GPU virtualization (nice to have).
  • Exposure to hardware validation, vendor collaboration, and long-term OEM roadmap alignment (nice to have).
  • Background in compute telemetry, real-time performance profiling, or large-scale A/B infrastructure testing (nice to have).
  • Legally authorized to work in the United States

Responsibilities

  • Architect and define scalable compute platforms optimized for AI/ML, simulation, and high-throughput workloads
  • Develop compute system standards and design patterns to ensure consistency, performance, and maintainability across infrastructure
  • Evaluate emerging CPU, GPU, and accelerator technologies, owning architectural tradeoff decisions affecting compute density, power, cooling, and total cost
  • Collaborate with product and engineering teams to map workload requirements to compute platform capabilities across bare metal and cloud deployments
  • Convert ambiguous business or customer needs into measurable platform requirements, technical specifications, acceptance criteria, and architecture decisions
  • Define compute platform roadmaps and architectural reference designs guiding hardware selection, firmware baselines, rack-level, and cluster design
  • Act as a technical lead during new platform introductions, guiding validation and performance characterization efforts
  • Mentor systems engineers and cross-functional stakeholders on compute performance tuning, sizing, and architectural decisions

Skills

GPU HPC
CPU Architecture
Performance Tuning
Resource Scheduling
Thermal Optimization
Power Optimization
Lifecycle Management
Architectural Tradeoffs
Validation & Testing
Telemetry
Real-Time Profiling
Influencing Strategy
Communication
Ownership
Self-Starter
Kubernetes
Slurm
NVLink
InfiniBand
RoCE
GPU Virtualization

Tools

NVLink
InfiniBand
RoCE
Slurm
Kubernetes
GPU Virtualization

Job description

  • Architect and define scalable compute platforms optimized for AI/ML, simulation, and high-throughput workloads
  • Develop compute system standards and design patterns to ensure consistency, performance, and maintainability across infrastructure
  • Evaluate emerging CPU, GPU, and accelerator technologies, owning architectural tradeoff decisions affecting compute density, power, cooling, and total cost
  • Collaborate with product and engineering teams to map workload requirements to compute platform capabilities across bare metal and cloud deployments
  • Convert ambiguous business or customer needs into measurable platform requirements, technical specifications, acceptance criteria, and architecture decisions
  • Define compute platform roadmaps and architectural reference designs guiding hardware selection, firmware baselines, rack-level, and cluster design
  • Act as a technical lead during new platform introductions, guiding validation and performance characterization efforts
  • Mentor systems engineers and cross-functional stakeholders on compute performance tuning, sizing, and architectural decisions
Requirements
  • Proven experience (7+ years) architecting large-scale 10k-100k+ GPU HPC or cloud compute platforms
  • Deep knowledge of CPU/GPU architectures, memory hierarchies, and accelerator topologies
  • Experience designing systems around high-bandwidth, low-latency fabrics (NVLink, InfiniBand, and RoCE)
  • Strong understanding of system performance tuning, resource scheduling, thermal and power optimization, and compute lifecycle management
  • Comfortable working across hardware and software boundaries, especially at the intersection of compute architecture, OS behavior, and orchestration layers
  • Skilled at balancing architectural tradeoffs for density, power efficiency, cooling, and performance
  • Strong analytical and communication skills, with a track record of influencing technical strategy across teams
  • Strong ownership and can do attitude, self-starter who feels comfortable working in ambiguity
  • Hands-on experience with AI/ML workloads and their compute performance characteristics (nice to have)
  • Familiarity with orchestration tools used in HPC, such as Slurm and Kubernetes (nice to have)
  • Experience with virtualization technologies, specifically GPU virtualization (nice to have)
  • Exposure to hardware validation, vendor collaboration, and long-term OEM roadmap alignment (nice to have)
  • Background in compute telemetry, real-time performance profiling, or large-scale A/B infrastructure testing (nice to have)
  • Legally authorized to work in the United States
Core Competencies

Demonstrates expertise in architecting large-scale GPU HPC and cloud compute platforms, with a strong focus on CPU/GPU architectures, performance tuning, and resource optimization. Capable of translating business needs into technical specifications and guiding cross-functional teams in the development of scalable compute solutions.

Highest-signal resume keywords
  • Architecting Large-Scale GPU HPC Platforms
  • CPU/GPU Architecture Knowledge
  • System Performance Tuning
  • High-Bandwidth, Low-Latency Fabric Design
  • AI/ML Workload Optimization
Hard Skills
  • Compute Platform Architecture
  • Performance Characterization
  • Resource Scheduling
  • Thermal Optimization
  • Power Optimization
  • Compute Lifecycle Management
  • Architectural Tradeoff Analysis
  • Validation and Testing
  • Compute Telemetry
  • Real-Time Performance Profiling
Soft Skills
  • Analytical Skills
  • Communication Skills
  • Influencing Technical Strategy
  • Ownership
  • Self-Starter
Industry Keywords
  • AI/ML
  • High-Throughput Workloads
  • Compute Density
  • Cloud Deployments
  • HPC
Tools & Technologies
  • NVLink
  • InfiniBand
  • RoCE
  • Slurm
  • Kubernetes
  • GPU Virtualization
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff HPC Network Architect
Staff HPC Network Architect

Jobtailor • California (MO)

On-site
USD 180,000 - 260,000
HPC Applications Engineer – Performance
HPC Applications Engineer – Performance

Jobtailor • Minnesota

On-site
USD 120,000 - 160,000
Senior System Engineer – GPU Platforms
Senior System Engineer – GPU Platforms

Jobtailor • San Jose (CA)

On-site
USD 150,000 - 210,000
HPC Infrastructure Engineer
HPC Infrastructure Engineer

Arcadia • San Francisco (CA)

On-site
USD 180,000 - 260,000
Member of Technical Staff - AI Infrastructure
Member of Technical Staff - AI Infrastructure

Veeda Innovation • California (MO)

Hybrid
USD 180,000 - 240,000
Sr HPC Hardware Engineer
Sr HPC Hardware Engineer

Career Techniques • Dallas (TX)

Hybrid
USD 120,000 - 180,000
GPU System Performance Architect
GPU System Performance Architect

Jobtailor • California (MO)

On-site
USD 180,000 - 260,000
Hardware Engineer, System Design
Hardware Engineer, System Design

Jobtailor • Mountain View (CA)

On-site
USD 180,000 - 240,000
VP – AI Infrastructure Engineering
VP – AI Infrastructure Engineering

Jobtailor • Bellevue (WA)

On-site
USD 200,000 - 350,000
Senior Solutions Engineer, AI Infrastructure
Senior Solutions Engineer, AI Infrastructure

VAST Data • New York (NY)

On-site
USD 150,000 - 200,000