HPC Systems Validation Lead - Cloud AI Infrastructure

Lambda Inc.

San Jose (CA)

Hybrid

USD 150,000 - 210,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Cash & equity compensation
Health, dental & vision coverage for您

Job summary

Lambda Inc. in San Jose is seeking an experienced Hardware Engineer to lead system integration validation across HPC/AI platforms, from NPI to deployment. You will drive functional, stress and reliability testing, coordinate across fleet, infrastructure, and PMO teams, and own lab setups for repeatable hardware evaluation.

The role requires deep expertise in hardware validation, performance benchmarking, and vendor engagement to ensure production-ready platforms for AI workloads.

Qualifications

  • 5+ years in hardware integration validation or related fields.
  • Experience with HPC, data centre, or cloud infra hardware.
  • Proficient in system integration testing and benchmarking at high levels.

Responsibilities

  • Own system integration validation for HPC/AI hardware during NPI and deployment.
  • Develop and execute functional, stress, reliability tests; enable scale testing.
  • Collaborate across teams to enable automation and large-scale benchmarking.
  • Manage firmware/software compatibility during NPI and post-production.
  • Support root-cause analysis and improve product reliability.
  • Set up lab/test infrastructure for repeatable hardware evaluation.
  • Engage with vendors for hardware validation and manufacturing testing.
  • Drive validation plans for key components (memory, SSDs, NICs, PSUs, etc).

Skills

Hardware integration
Fleet reliability engineering
Performance benchmarking
NPI cycles
Vendor management
Lab/test automation
Root cause analysis

Education

Bachelor's degree in Electrical Engineering

Tools

BMC
BIOS configurations
NIC configurations

Job description

Lambda Inc. in San Jose is seeking an experienced Hardware Engineer to lead system integration validation across HPC/AI platforms, from NPI to deployment. You will drive functional, stress and reliability testing, coordinate across fleet, infrastructure, and PMO teams, and own lab setups for repeatable hardware evaluation.

The role requires deep expertise in hardware validation, performance benchmarking, and vendor engagement to ensure production-ready platforms for AI workloads.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior HPC Systems Validation Engineer – AI Cloud Infra
Senior HPC Systems Validation Engineer – AI Cloud Infra

Neura Market • San Jose (CA)

Hybrid
USD 180,000 - 230,000
Senior HPC Validation Engineer – AI Cloud Infrastructure
Senior HPC Validation Engineer – AI Cloud Infrastructure

Lambda • San Jose (CA)

On-site
USD 150,000 - 210,000
Health coverage
Dental coverage
Vision coverage
+4
Hands-on HPC Platform Hardware Lead for AI Cloud
Hands-on HPC Platform Hardware Lead for AI Cloud

Lambda Inc. • San Jose (CA)

Hybrid
USD 180,000 - 280,000
Health, dental, and vision coverage
Wellness stipend
Commuter stipend
+2
Senior HPC Hardware Platform Lead for AI Cloud
Senior HPC Hardware Platform Lead for AI Cloud

Lambda Labs • United States

On-site
USD 150,000 - 230,000
Health insurance
Dental and vision coverage
Wellness stipend
+3
Senior HPC Platform Hardware Engineer - Hybrid (San Jose)
Senior HPC Platform Hardware Engineer - Hybrid (San Jose)

Neura Market • San Jose (CA)

On-site
USD 180,000 - 280,000
Cash & equity compensation
Health, dental, and vision coverage
Wellness stipends
+1
Senior HPC Systems Validation Engineer
Senior HPC Systems Validation Engineer

Lambda Inc. • San Jose (CA)

Hybrid
USD 150,000 - 210,000
Cash & equity compensation
Health, dental & vision coverage for您
Senior HPC Systems Validation Engineer
Senior HPC Systems Validation Engineer

Lambda • San Jose (CA)

On-site
USD 150,000 - 210,000
Health coverage
Dental coverage
Vision coverage
+4
Senior AI Cloud SRE — HPC & GPU Reliability Lead
Senior AI Cloud SRE — HPC & GPU Reliability Lead

Lambda • San Francisco (CA)

On-site
USD 180,000 - 230,000
Health, dental, and vision
401k with company match
Flexible paid time off
+2
Senior HPC Systems Validation Engineer
Senior HPC Systems Validation Engineer

Neura Market • San Jose (CA)

Hybrid
USD 180,000 - 230,000
Senior HPC Systems Architect – Liquid-Cooled AI GPU Infra
Senior HPC Systems Architect – Liquid-Cooled AI GPU Infra

Lambda Labs • United States

Hybrid
USD 180,000 - 280,000
Equity compensation
Health, dental and vision coverage
401k with 2% company match
+1