Senior HPC Validation Engineer – AI Cloud Infrastructure

Lambda

San Jose (CA)

On-site

USD 150,000 - 210,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Health coverage
Dental coverage
Vision coverage
Wellness stipend
Commuter stipend
401k with match
Flexible PTO

Job summary

Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure based in San Jose. This role focuses on validating and integrating server hardware for HPC AI/ML deployments, driving NPI readiness, and collaborating across engineering, supply chain, and operations to deliver production-grade hardware solutions.

You will own hands-on lab setups, run comprehensive tests, and coordinate with vendors to qualify components while supporting RMA and reliability initiatives.

Qualifications

  • 5 years of experience in hardware integration validation or related roles.
  • Experience with fleet hardware reliability engineering and component qualification.
  • Knowledge of performance benchmarking for HPC, data center, or cloud infrastructure.
  • Experience with vendor-led NPI cycles and production support.

Responsibilities

  • Own system integration validation for new HPC AI/ML hardware platforms throughout NPI and deployment readiness.
  • Develop and execute functional, stress, reliability, and benchmark tests on new hardware systems.
  • Collaborate with multiple engineering and PMO teams to enable tooling automation and benchmarking.
  • Manage firmware and software compatibility, and advise on releases during NPI and production.
  • Support RMA triage and root-cause analysis to improve product reliability.
  • Set up labs and test infrastructure for repeatable hardware evaluation and debugging.

Skills

Hardware integration validation
Fleet hardware reliability
Component qualification
Performance validation
NPI coordination

Tools

PLM systems
BOM structure
BIOS settings
Firmware testing

Job description

Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure based in San Jose. This role focuses on validating and integrating server hardware for HPC AI/ML deployments, driving NPI readiness, and collaborating across engineering, supply chain, and operations to deliver production-grade hardware solutions.

You will own hands-on lab setups, run comprehensive tests, and coordinate with vendors to qualify components while supporting RMA and reliability initiatives.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior HPC Systems Validation Engineer – AI Cloud Infra
Senior HPC Systems Validation Engineer – AI Cloud Infra

Neura Market • San Jose (CA)

Hybrid
USD 180,000 - 230,000
HPC Systems Validation Lead - Cloud AI Infrastructure
HPC Systems Validation Lead - Cloud AI Infrastructure

Lambda Inc. • San Jose (CA)

Hybrid
USD 150,000 - 210,000
Cash & equity compensation
Health, dental & vision coverage for您
Senior HPC Platform Hardware Engineer - Hybrid (San Jose)
Senior HPC Platform Hardware Engineer - Hybrid (San Jose)

Neura Market • San Jose (CA)

On-site
USD 180,000 - 280,000
Cash & equity compensation
Health, dental, and vision coverage
Wellness stipends
+1
Senior HPC Hardware Platform Lead for AI Cloud
Senior HPC Hardware Platform Lead for AI Cloud

Lambda Labs • United States

On-site
USD 150,000 - 230,000
Health insurance
Dental and vision coverage
Wellness stipend
+3
Hands-on HPC Platform Hardware Lead for AI Cloud
Hands-on HPC Platform Hardware Lead for AI Cloud

Lambda Inc. • San Jose (CA)

Hybrid
USD 180,000 - 280,000
Health, dental, and vision coverage
Wellness stipend
Commuter stipend
+2
Senior HPC Systems Validation Engineer
Senior HPC Systems Validation Engineer

Lambda Inc. • San Jose (CA)

Hybrid
USD 150,000 - 210,000
Cash & equity compensation
Health, dental & vision coverage for您
Senior HPC Architect: GPU Clusters & Liquid Cooling
Senior HPC Architect: GPU Clusters & Liquid Cooling

Neura Market • San Jose (CA)

Hybrid
USD 180,000 - 240,000
Health, dental, and vision
401k with company match
Wellness stipend
+1
Senior AI Cloud SRE — HPC & GPU Reliability Lead
Senior AI Cloud SRE — HPC & GPU Reliability Lead

Lambda • San Francisco (CA)

On-site
USD 180,000 - 230,000
Health, dental, and vision
401k with company match
Flexible paid time off
+2
Senior AI Cloud SRE – HPC & GPU Infra (Hybrid)
Senior AI Cloud SRE – HPC & GPU Infra (Hybrid)

Lambda • United States

Hybrid
USD 140,000 - 200,000
Senior HPC Systems Validation Engineer
Senior HPC Systems Validation Engineer

Neura Market • San Jose (CA)

Hybrid
USD 180,000 - 230,000