Senior HPC Systems Architect: Liquid-Cooled GPU Cloud

Socket.dev

San Jose (CA)

Hybrid

USD 180,000 - 260,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Lambda, The Superintelligence Cloud, is seeking a Senior HPC Systems Architect to design and lead the development of large-scale GPU clusters and HPC infrastructure. This role requires presence in the San Francisco or San Jose office 4 days per week, with Tuesday as the designated work-from-home day.

You will define system requirements, build testing frameworks, evaluate new technologies, and mentor engineering teams to deliver scalable, production-ready solutions for AI workloads.

Qualifications

  • 8+ years of experience designing and architecting large-scale HPC and distributed computing systems.
  • Expert-level knowledge of HPC hardware including GPU clusters, compute nodes, InfiniBand and Ethernet, and distributed storage.
  • Hands-on experience with direct-to-chip liquid cooling systems.
  • Proven expertise in creating robust performance benchmarks, capacity planning, and system validation.
  • Exceptional skills in system architecture, design documentation, and technical specifications.
  • Ability to work collaboratively across teams, ensuring alignment of technical solutions with business objectives.

Responsibilities

  • Design and architect advanced HPC systems optimized for large-scale computational workloads and AI applications.
  • Collaborate with internal teams and stakeholders to define system requirements and performance goals.
  • Develop comprehensive testing frameworks to rigorously assess system performance, scalability, and reliability.
  • Evaluate emerging technologies and architectural approaches to continuously enhance infrastructure capabilities.
  • Create detailed architectural plans, documentation, and blueprints to guide implementation teams.
  • Provide technical leadership and mentoring to engineering teams, fostering best practices in HPC architecture.

Skills

HPC architecture
GPU clusters
High-speed networking
System benchmarking
Documentation
Cross-team collaboration
Problem solving

Tools

Ansible
Terraform
Kubernetes
InfiniBand

Job description

Lambda, The Superintelligence Cloud, is seeking a Senior HPC Systems Architect to design and lead the development of large-scale GPU clusters and HPC infrastructure. This role requires presence in the San Francisco or San Jose office 4 days per week, with Tuesday as the designated work-from-home day.

You will define system requirements, build testing frameworks, evaluate new technologies, and mentor engineering teams to deliver scalable, production-ready solutions for AI workloads.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior HPC Architect: GPU Clusters & Liquid Cooling
Senior HPC Architect: GPU Clusters & Liquid Cooling

Neura Market • San Jose (CA)

Hybrid
USD 180,000 - 240,000
Health, dental, and vision
401k with company match
Wellness stipend
+1
Senior HPC Systems Architect – Liquid-Cooled AI GPU Infra
Senior HPC Systems Architect – Liquid-Cooled AI GPU Infra

Lambda Labs • United States

Hybrid
USD 180,000 - 280,000
Equity compensation
Health, dental and vision coverage
401k with 2% company match
+1
Senior HPC Systems Architect: Liquid-Cooled GPU AI Infra
Senior HPC Systems Architect: Liquid-Cooled GPU AI Infra

Lambda • San Jose (CA)

Hybrid
USD 180,000 - 280,000
401k Plan
Health, dental, and vision coverage
Wellness stipend
+2
Hybrid HPC Systems Architect - GPU Cloud for AI
Hybrid HPC Systems Architect - GPU Cloud for AI

The Consensus • San Jose (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Cash compensation
Equity compensation
Health, dental and vision coverage
+1
Senior Cloud Platform Engineer - GPU Infrastructure
Senior Cloud Platform Engineer - GPU Infrastructure

Socket.dev • San Francisco (CA)

On-site
USD 180,000 - 240,000
Health, dental, and vision coverage
Wellness stipend
Commuter stipend
+1
Senior HPC Systems Architect
Senior HPC Systems Architect

Socket.dev • San Jose (CA)

Hybrid
USD 180,000 - 260,000
Senior HPC Systems Architect
Senior HPC Systems Architect

The Consensus • San Jose (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Cash compensation
Equity compensation
Health, dental and vision coverage
+1
Senior HPC Systems Architect
Senior HPC Systems Architect

Lambda • San Francisco (CA)

Hybrid
USD 180,000 - 240,000
HPC Systems Architect
HPC Systems Architect

Lambda • San Jose (CA)

Hybrid
USD 180,000 - 280,000
401k Plan
Health, dental, and vision coverage
Wellness stipend
+2
Senior HPC Systems Architect
Senior HPC Systems Architect

Lambda Labs • United States

Hybrid
USD 180,000 - 280,000
Equity compensation
Health, dental and vision coverage
401k with 2% company match
+1