Senior Cloud Infrastructure Engineer – GPU & DPU

Lambda

United States

Remote

USD 180,000 - 260,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Lambda, The Superintelligence Cloud, is seeking an experienced engineer to design, implement, and maintain production systems powering GPU fleet lifecycle management at scale. You will automate provisioning, configuration, and deployment and support hardware introductions for new server platforms.

You will also improve lifecycle workflows, debug firmware, BIOS, and networking issues, and collaborate across infrastructure, security and product teams to deliver scalable solutions.

Qualifications

  • 6+ years of production experience with Go or Python in large-scale systems.
  • Deep Linux OS, hardware, and networking debugging skills.
  • Hands-on with bare metal provisioning and lifecycle tooling.
  • Experience with Redfish, BMC, IPMI, DHCP, and PXE preferred.
  • Ability to coordinate across software, infrastructure and vendor teams.

Responsibilities

  • Develop and maintain production systems for GPU fleet lifecycle management.
  • Automate infrastructure, provisioning, and deployment workflows.
  • Support hardware introduction for new server and accelerator platforms.
  • Improve machine lifecycle processes and firmware update workflows.
  • Assist in DPU lifecycle automation and debugging of hardware issues.
  • Collaborate with infrastructure, security, and product engineers.

Skills

Go (Golang)
Python
Bare metal hardware management
Linux debugging
Cross-team collaboration
DPUs / NICs

Tools

Redfish
BMC
IPMI
DHCP
PXE

Job description

Lambda, The Superintelligence Cloud, is seeking an experienced engineer to design, implement, and maintain production systems powering GPU fleet lifecycle management at scale. You will automate provisioning, configuration, and deployment and support hardware introductions for new server platforms.

You will also improve lifecycle workflows, debug firmware, BIOS, and networking issues, and collaborate across infrastructure, security and product teams to deliver scalable solutions.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Cloud Platform Engineer - GPU Infra, Hybrid
Senior Cloud Platform Engineer - GPU Infra, Hybrid

Neura Market • San Francisco (CA)

Hybrid
USD 180,000 - 240,000
Wellness stipend
Commuter stipend
401k with company match
Senior Cloud Platform Engineer, GPU Core & Lifecycle
Senior Cloud Platform Engineer, GPU Core & Lifecycle

Lambda Labs • United States

Hybrid
USD 180,000 - 240,000
Health insurance
Dental insurance
Vision insurance
+5
Senior Cloud Platform Engineer - GPU Infrastructure
Senior Cloud Platform Engineer - GPU Infrastructure

Socket.dev • San Francisco (CA)

On-site
USD 180,000 - 240,000
Health, dental, and vision coverage
Wellness stipend
Commuter stipend
+1
Staff Cloud Compute Engineer – GPU/CPU Lifecycle
Staff Cloud Compute Engineer – GPU/CPU Lifecycle

Neura Market • Bellevue (WA), Northern (KY)

Hybrid
USD 190,000 - 240,000
Cash & equity compensation
Health, dental, vision coverage
Wellness stipend
+2
Senior Cloud Platform Engineer — GPU Compute Orchestration
Senior Cloud Platform Engineer — GPU Compute Orchestration

Lambda Inc. • San Francisco (CA)

Hybrid
USD 180,000 - 240,000
Health, dental, and vision coverage
401(k) plan with company match (USA)
Data Center Systems Engineer for GPU Infrastructure
Data Center Systems Engineer for GPU Infrastructure

Lambda • Quincy (WA)

On-site
USD 70,000 - 110,000
Health insurance
Dental coverage
Vision coverage
+2
Staff Compute Platform Engineer – AI Cloud & GPU/CPU
Staff Compute Platform Engineer – AI Cloud & GPU/CPU

Lambda • San Francisco (CA)

On-site
USD 180,000 - 240,000
Health, dental, vision coverage
401k with 2% company match
Flexible PTO
+1
Staff Software Engineer — AI Cloud Compute Leader
Staff Software Engineer — AI Cloud Compute Leader

AI Chopping Block • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Wellness stipend
Commuter stipend
401k Plan with company match
Staff Compute Cloud Engineer - GPU-First Lifecycle & Control
Staff Compute Cloud Engineer - GPU-First Lifecycle & Control

Lambda • San Jose (CA)

On-site
USD 314,000 - 465,000
Health, dental, and vision coverage
401k Plan with 2% company match
Flexible paid time off
+1
Staff Software Engineer, Cloud Compute & GPU Lifecycle
Staff Software Engineer, Cloud Compute & GPU Lifecycle

Lambda • United States

Remote
USD 180,000 - 320,000