Hybrid HPC Systems Architect - GPU Cloud for AI

The Consensus

San Jose, Northern (CA, KY)

Hybrid

USD 180,000 - 240,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Cash compensation
Equity compensation
Health, dental and vision coverage
Wellness stipend

Job summary

Lambda is seeking a Senior HPC Systems Architect to lead the design, development, and testing of large-scale, liquid-cooled HPC infrastructures for AI workloads. This role demands deep expertise in GPU clusters, high-performance networking, and scalable storage, with a proven track record of architectural leadership.

The position involves guiding cross-functional teams, creating detailed architectural plans, and mentoring engineers to uphold best practices in HPC architecture.

Qualifications

  • 8+ years of experience designing and architecting large-scale HPC and distributed computing systems.
  • Expert-level knowledge of HPC hardware including GPU clusters, compute nodes, high-speed networking (InfiniBand, Ethernet), and distributed storage.
  • Hands-on experience with direct-to-chip liquid cooling systems.
  • Proven expertise in creating robust performance benchmarks, capacity planning, and system validation.
  • Exceptional skills in system architecture, design documentation, and technical specifications.
  • Ability to work collaboratively across teams, ensuring alignment of technical solutions with business objectives.
  • Self-motivated, strategic thinker with strong analytical and problem-solving capabilities.

Responsibilities

  • Design and architect advanced HPC systems optimized for large-scale computational workloads and AI applications.
  • Collaborate with internal teams and stakeholders to define system requirements and performance goals.
  • Develop comprehensive testing frameworks to rigorously assess system performance, scalability, and reliability.
  • Evaluate emerging technologies and architectural approaches to continuously enhance infrastructure capabilities.
  • Create detailed architectural plans, documentation, and blueprints to guide implementation teams.
  • Provide technical leadership and mentoring to engineering teams, fostering best practices in HPC architecture.

Skills

HPC systems design
GPU clusters
High-speed networking
Liquid cooling
System validation
Technical leadership

Tools

Ansible
Terraform
Kubernetes

Job description

Lambda is seeking a Senior HPC Systems Architect to lead the design, development, and testing of large-scale, liquid-cooled HPC infrastructures for AI workloads. This role demands deep expertise in GPU clusters, high-performance networking, and scalable storage, with a proven track record of architectural leadership.

The position involves guiding cross-functional teams, creating detailed architectural plans, and mentoring engineers to uphold best practices in HPC architecture.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior HPC Systems Architect – Liquid-Cooled AI GPU Infra
Senior HPC Systems Architect – Liquid-Cooled AI GPU Infra

Lambda Labs • United States

Hybrid
USD 180,000 - 280,000
Equity compensation
Health, dental and vision coverage
401k with 2% company match
+1
Senior HPC Systems Architect: Liquid-Cooled GPU Cloud
Senior HPC Systems Architect: Liquid-Cooled GPU Cloud

Socket.dev • San Jose (CA)

Hybrid
USD 180,000 - 260,000
Senior HPC Architect: GPU Clusters & Liquid Cooling
Senior HPC Architect: GPU Clusters & Liquid Cooling

Neura Market • San Jose (CA)

Hybrid
USD 180,000 - 240,000
Health, dental, and vision
401k with company match
Wellness stipend
+1
GPU-First Cloud Compute Engineer (Hybrid)
GPU-First Cloud Compute Engineer (Hybrid)

Lambda • San Francisco (CA)

Hybrid
USD 266,000 - 395,000
Health, dental, and vision coverage
Wellness stipend
Commuter stipend
+1
Senior Cloud Platform Engineer - GPU Infra, Hybrid
Senior Cloud Platform Engineer - GPU Infra, Hybrid

Neura Market • San Francisco (CA)

Hybrid
USD 180,000 - 240,000
Wellness stipend
Commuter stipend
401k with company match
Senior HPC Systems Architect
Senior HPC Systems Architect

The Consensus • San Jose (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Cash compensation
Equity compensation
Health, dental and vision coverage
+1
Senior HPC Systems Architect
Senior HPC Systems Architect

Socket.dev • San Jose (CA)

Hybrid
USD 180,000 - 260,000
Senior HPC Systems Architect
Senior HPC Systems Architect

Lambda Labs • United States

Hybrid
USD 180,000 - 280,000
Equity compensation
Health, dental and vision coverage
401k with 2% company match
+1
Senior HPC Systems Architect
Senior HPC Systems Architect

Lambda • San Francisco (CA)

Hybrid
USD 180,000 - 240,000
Senior HPC Systems Architect
Senior HPC Systems Architect

Neura Market • San Jose (CA)

Hybrid
USD 180,000 - 240,000
Health, dental, and vision
401k with company match
Wellness stipend
+1