Senior HPC Compute Architect for AI Cloud

Socket.dev

San Jose (CA)

On-site

USD 180,000 - 240,000

Full time

3 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of thousands of customers. Our customers range from AI researchers to enterprises and hyperscalers. Lambda's mission is to make compute as ubiquitous as electricity and give everyone the power of superintelligence.

One person, one GPU. If you'd like to build the world's best AI cloud, join us. This position requires presence in our San Jose or San Francisco office location 4 days per week; Lambda’s

Qualifications

  • 7+ years of experience architecting large-scale GPU HPC or cloud compute platforms.
  • Deep knowledge of CPU/GPU architectures, memory hierarchies and accelerator topologies.
  • Experience with high-bandwidth, low-latency fabrics (NVLink, InfiniBand, RoCE).
  • Strong performance tuning, resource scheduling, thermal and power optimization, and compute lifecycle management.

Responsibilities

  • Architect and define scalable compute platforms optimized for AI/ML, simulation and high-throughput workloads.
  • Develop compute system standards and design patterns for consistency, performance and maintainability.
  • Evaluate CPU, GPU and accelerator technologies and make trade-off decisions for density, power and cost.
  • Collaborate with product and engineering to map workloads to compute capabilities across bare metal and cloud deployments.
  • Lead platform introductions, guiding validation and performance characterization.
  • Mentor engineers on compute performance tuning and architectural decisions.

Skills

GPU HPC
CPU/GPU architectures
high-bandwidth fabrics
Performance tuning
Compute lifecycle management
OS & orchestration
NVLink InfiniBand RoCE

Tools

Slurm
Kubernetes
GPU virtualization

Job description

Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of thousands of customers. Our customers range from AI researchers to enterprises and hyperscalers. Lambda's mission is to make compute as ubiquitous as electricity and give everyone the power of superintelligence.

One person, one GPU. If you'd like to build the world's best AI cloud, join us. This position requires presence in our San Jose or San Francisco office location 4 days per week; Lambda’s

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior HPC Architect: GPU Clusters & Liquid Cooling
Senior HPC Architect: GPU Clusters & Liquid Cooling

Neura Market • San Jose (CA)

Hybrid
USD 180,000 - 240,000
Health, dental, and vision
401k with company match
Wellness stipend
+1
Senior HPC Architect for AI Compute Platforms
Senior HPC Architect for AI Compute Platforms

Lambda • San Jose (CA)

On-site
USD 180,000 - 260,000
Health, dental, and vision coverage
Equity compensation
401k with 2% company match
+3
Senior HPC Validation Engineer – AI Cloud Infrastructure
Senior HPC Validation Engineer – AI Cloud Infrastructure

Lambda • San Jose (CA)

On-site
USD 150,000 - 210,000
Health coverage
Dental coverage
Vision coverage
+4
Staff HPC Systems Architect
Staff HPC Systems Architect

Lambda • San Jose (CA)

On-site
USD 180,000 - 260,000
Health, dental, and vision coverage
Equity compensation
401k with 2% company match
+3
Senior HPC Systems Architect
Senior HPC Systems Architect

Lambda • San Francisco (CA)

Hybrid
USD 180,000 - 240,000
Senior HPC Systems Architect
Senior HPC Systems Architect

Neura Market • San Jose (CA)

Hybrid
USD 180,000 - 240,000
Health, dental, and vision
401k with company match
Wellness stipend
+1
Senior HPC Systems Architect
Senior HPC Systems Architect

The Consensus • San Jose (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Cash compensation
Equity compensation
Health, dental and vision coverage
+1
Senior HPC Systems Architect
Senior HPC Systems Architect

Socket.dev • San Jose (CA)

Hybrid
USD 180,000 - 260,000
Lead HPC Network Architect for AI Cloud
Lead HPC Network Architect for AI Cloud

Socket.dev • San Jose (CA)

Hybrid
USD 180,000 - 260,000
Health, dental, and vision coverage
Equity compensation
401k with company match
+2
AI Operations Engineer - IT/Internal Infrastructure
AI Operations Engineer - IT/Internal Infrastructure

Lambda • San Jose (CA)

On-site
USD 206,000 - 275,000