Senior SRE: AI Cloud Platform & Kubernetes Expert

Lambda

Bellevue (WA)

On-site

USD 180,000 - 260,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Health insurance
Dental insurance
Vision insurance
Wellness stipend
401k with company match
Flexible paid time off

Job summary

Lambda seeks a Senior Site Reliability Engineer to scale and harden its AI cloud platform across data centers. You’ll improve provisioning reliability, implement robust monitoring, and drive incident response and postmortems.

You will define SLIs/SLOs, automate drift remediation, and build disaster recovery workflows using Terraform, Argo CD, and Kubernetes-native tools. Collaboration with multiple teams and mentorship are key aspects of the role.

Qualifications

  • 7+ years of experience in site reliability, infrastructure, distributed systems, or production software engineering.
  • Deep experience operating Kubernetes in production.

Responsibilities

  • Operate and scale critical platform services across Lambda’s data centers.
  • Improve the reliability of compute provisioning, Instance lifecycle, and regional orchestration systems.
  • Build monitoring, alerting, and tracing for service health, provisioning latency, and customer-impacting failures.
  • Define SLIs, SLOs, error budgets, and operational readiness standards.
  • Automate detection and remediation of configuration drift, failed workflows, and orphaned resources.
  • Build safe deployment, rollback, and disaster recovery workflows using infrastructure as code and GitOps.
  • Design fault-isolation mechanisms that reduce blast radius and prevent cascading failures.
  • Lead production incident response, postmortems, and durable corrective actions.
  • Partner with Compute, Networking, Storage, Security, and Support teams.
  • Participate in on-call and improve its sustainability through automation and better tooling.
  • Mentor engineers and raise the reliability bar across the organization.

Skills

Kubernetes production experience
Infrastructure as Code
CI/CD / GitOps
Observability stack
Go or Python
SLIs / SLOs
Incident response leadership
Cross-team communication

Tools

Terraform
Argo CD
Flux
Helm
Kustomize
OpenTelemetry
Prometheus
Grafana
Datadog

Job description

Lambda seeks a Senior Site Reliability Engineer to scale and harden its AI cloud platform across data centers. You’ll improve provisioning reliability, implement robust monitoring, and drive incident response and postmortems.

You will define SLIs/SLOs, automate drift remediation, and build disaster recovery workflows using Terraform, Argo CD, and Kubernetes-native tools. Collaboration with multiple teams and mentorship are key aspects of the role.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior SRE – AI Cloud Platform, Kubernetes Expert
Senior SRE – AI Cloud Platform, Kubernetes Expert

Socket.dev • San Francisco (CA)

On-site
USD 180,000 - 240,000
Health, dental, vision coverage for in
Wellness and commuter stipends
401k with 2% company match
+1
Senior SRE: AI Cloud Platform & Kubernetes
Senior SRE: AI Cloud Platform & Kubernetes

Lambda Inc. • San Francisco (CA)

Hybrid
USD 190,000 - 270,000
Health insurance
401k with company match
Flexible PTO
+2
Senior SRE: Managed Kubernetes for AI Cloud
Senior SRE: Managed Kubernetes for AI Cloud

Lambda Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Health insurance
Dental insurance
Vision insurance
+3
Senior SRE — Core Cloud Platform Resilience & Automation
Senior SRE — Core Cloud Platform Resilience & Automation

Front Door Defense • San Jose (CA)

On-site
USD 240,000 - 356,000
Health, dental, and vision coverage
401k Plan with company match
Flexible paid time off
+1
Senior SRE - Managed Kubernetes for AI Cloud
Senior SRE - Managed Kubernetes for AI Cloud

Socket.dev • San Francisco (CA)

Hybrid
USD 150,000 - 230,000
Health, dental, vision
4-day in-office work week
Wellness stipend
+1
Senior SRE: Managed Kubernetes for AI Cloud Platforms
Senior SRE: Managed Kubernetes for AI Cloud Platforms

Lambda • Bellevue (WA)

On-site
USD 240,000 - 356,000
Health, dental, and vision coverage
401k matching
Flexible PTO
+1
Senior SRE: Bare-Metal Kubernetes for AI Cloud
Senior SRE: Bare-Metal Kubernetes for AI Cloud

AI Chopping Block • San Francisco (CA), Northern (KY)

Hybrid
USD 190,000 - 270,000
Health, dental, vision
401k with company match
Wellness stipends
+1
Senior AI Cloud SRE — HPC & GPU Infra
Senior AI Cloud SRE — HPC & GPU Infra

Lambda Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Health insurance
Dental insurance
Vision insurance
+4
Senior SRE: Managed Kubernetes for AI Cloud (Hybrid)
Senior SRE: Managed Kubernetes for AI Cloud (Hybrid)

Lambda • San Francisco (CA)

Hybrid
USD 170,000 - 260,000
401k Plan with company match (USA)
Health, dental, and vision coverage
Wellness and commuter stipends
+1
Senior AI Cloud SRE — Hybrid (SF/Bellevue)
Senior AI Cloud SRE — Hybrid (SF/Bellevue)

AI Chopping Block • San Francisco (CA), Northern (KY)

Hybrid
USD 170,000 - 250,000
Health, dental, vision coverage
Wellness stipend
Commuter stipend
+2