Senior Cloud Platform Engineer for AI GPU Infra

Lambda

San Jose (CA)

On-site

USD 170,000 - 260,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Health, dental, and vision coverage
Wellness and commuter stipends
401k Plan with company match
Flexible paid time off

Job summary

Lambda is seeking a Senior Software Engineer for its Core Cloud Platform in San Jose/San Francisco. You will build the control-plane systems powering Lambda’s GPU cloud, covering compute lifecycle, bare metal orchestration, maintenance actions, deployment readiness, reliability, and operational tooling.

You will work on APIs, workflows, state machines, and orchestration services, and collaborate with infrastructure, networking, security, and product teams to deliver end-to-end cloud

Qualifications

  • Bachelor’s degree or equivalent working experience.
  • 6+ years of professional software engineering experience building production backend or distributed systems.
  • Strong in Python, Go, or a similar backend/system language.
  • Experience designing and operating APIs, workflow engines, schedulers, orchestration services, or other distributed systems.
  • Understand reliability fundamentals: fault tolerance, idempotency, retries, state machines, failure handling, and production debugging.
  • Experience with cloud or cloud-like infrastructure primitives such as compute, networking, storage, capacity management, identity, or fleet operations.
  • Comfortable with Linux, containers, Kubernetes, infra automation, and service deployment patterns.
  • Have owned production services, participated in on-call, and improved systems based on operational learnings.
  • Care about testability, CI/CD, observability, metrics, logging, alerting, and supportable operations.
  • Can take ambiguous infrastructure problems and drive them to clear designs, implementation plans, and production outcomes.
  • Communicate clearly across engineering, product, support, infrastructure teams, and leadership.

Responsibilities

  • Build and operate core cloud platform services for compute lifecycle, bare metal hosts, capacity, placement, and maintenance workflows.
  • Design reliable APIs, backend services, state machines, and orchestration systems that power Lambda’s GPU cloud.
  • Work on bare metal lifecycle systems including launch, terminate, restart/reboot, host reclaim, validation, quarantine, and return-to-pool workflows.
  • Improve deployment, observability, testing, alerting, runbooks, and operational readiness for business-critical control-plane services.
  • Debug complex production issues across distributed services, infrastructure dependencies, networking, and cloud workflows.
  • Partner with infrastructure, networking, fleet, security, support, and product teams to define cross-system contracts and deliver end-to-end cloud capabilities.
  • Contribute to architecture, design docs, code reviews, incident follow-through, and mentoring across the team.

Skills

Python
Go
Distributed systems
Backend services

Education

Bachelor's degree or equivalent

Tools

Linux
Kubernetes
CI/CD

Job description

Lambda is seeking a Senior Software Engineer for its Core Cloud Platform in San Jose/San Francisco. You will build the control-plane systems powering Lambda’s GPU cloud, covering compute lifecycle, bare metal orchestration, maintenance actions, deployment readiness, reliability, and operational tooling.

You will work on APIs, workflows, state machines, and orchestration services, and collaborate with infrastructure, networking, security, and product teams to deliver end-to-end cloud

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Cloud Platform Engineer - GPU Infrastructure
Senior Cloud Platform Engineer - GPU Infrastructure

Socket.dev • San Francisco (CA)

On-site
USD 180,000 - 240,000
Health, dental, and vision coverage
Wellness stipend
Commuter stipend
+1
Senior Cloud Platform Engineer, GPU Core & Lifecycle
Senior Cloud Platform Engineer, GPU Core & Lifecycle

Lambda Labs • United States

Hybrid
USD 180,000 - 240,000
Health insurance
Dental insurance
Vision insurance
+5
Senior Cloud Platform Engineer - GPU Infra, Hybrid
Senior Cloud Platform Engineer - GPU Infra, Hybrid

Neura Market • San Francisco (CA)

Hybrid
USD 180,000 - 240,000
Wellness stipend
Commuter stipend
401k with company match
Senior Cloud Platform Engineer: Core GPU Infrastructure
Senior Cloud Platform Engineer: Core GPU Infrastructure

Lambda • San Francisco (CA)

Hybrid
USD 296,000 - 346,000
Health insurance
Dental insurance
Vision insurance
+2
Senior Cloud Platform Engineer — GPU Compute Orchestration
Senior Cloud Platform Engineer — GPU Compute Orchestration

Lambda Inc. • San Francisco (CA)

Hybrid
USD 180,000 - 240,000
Health, dental, and vision coverage
401(k) plan with company match (USA)
Staff Software Engineer — AI Cloud Compute Platform
Staff Software Engineer — AI Cloud Compute Platform

Lambda Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 300,000
Health, dental, vision coverage
Wellness stipend
401k with 2% company match
+2
Senior AI Cloud Solutions Engineer
Senior AI Cloud Solutions Engineer

Neura Market • San Francisco (CA)

Hybrid
USD 180,000 - 240,000
Health coverage
Dental coverage
Vision coverage
+3
Staff Compute Platform Engineer – AI Cloud (Remote)
Staff Compute Platform Engineer – AI Cloud (Remote)

Applied Methods Ltd • Bellevue (WA), Northern (KY)

Hybrid
USD 190,000 - 260,000
Health coverage
Dental coverage
Vision coverage
+2
Senior Cloud Infrastructure Engineer – GPU & DPU
Senior Cloud Infrastructure Engineer – GPU & DPU

Lambda • United States

Remote
USD 180,000 - 260,000
Senior GPU Cloud Solutions Engineer
Senior GPU Cloud Solutions Engineer

Lambda • San Francisco (CA)

Hybrid
USD 170,000 - 210,000
Health, dental, and vision coverage
401k with 2% company match
Wellness and commuter stipends
+1