Senior SRE – AI Cloud Platform, Kubernetes Expert

Socket.dev

San Francisco (CA)

On-site

USD 180,000 - 240,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Health, dental, vision coverage for in
Wellness and commuter stipends
401k with 2% company match
Flexible paid time off

Job summary

Lambda’s Core Cloud Platform seeks a Senior Site Reliability Engineer to boost reliability and scalability of compute provisioning and orchestration across data centers. You will work on Kubernetes, automation, observability, and incident response to lay a foundation for customer AI workloads.

Responsibilities include building monitoring and SLOs, automating remediation, and leading high-severity incidents.

Qualifications

  • 7+ years of site reliability, infrastructure, distributed systems, or production software engineering.
  • Deep experience operating Kubernetes in production.
  • Understand Kubernetes architecture, scheduling, networking, resource management, upgrades, and common failure modes.
  • Experience with physical data centers, private cloud, hybrid cloud, or environments without full reliance on managed services.
  • Proficient with Terraform or similar infrastructure-as-code tools.
  • CI/CD or GitOps workflows using Argo CD, Flux, Helm, or Kustomize.
  • Observability platforms such as OpenTelemetry, Prometheus, Grafana, or Datadog.
  • Production-quality tooling in Go or Python.

Responsibilities

  • Operate and scale critical platform services across Lambda’s data centers.
  • Improve reliability of compute provisioning, instance lifecycle, and regional orchestration systems.
  • Build monitoring, alerting, and tracing for service health and provisioning latency.
  • Define SLIs, SLOs, error budgets, and operational readiness standards.
  • Automate detection/remediation of configuration drift and failed workflows.
  • Build safe deployment, rollback, and disaster recovery workflows using IaC and GitOps.
  • Lead production incident response, postmortems, and durable corrective actions.
  • Collaborate with Compute, Networking, Storage, Security, and Support teams.

Skills

7+ years experience
Kubernetes in production
Kubernetes architecture
Terraform or IaC
CI/CD or GitOps
OpenTelemetry/Prometheus/Grafana/Datad
Go/Python tooling
SLIs/SLOs
incident leadership
communication across teams
ownership

Tools

Terraform
Argo CD
Flux
Helm
Kustomize

Job description

Lambda’s Core Cloud Platform seeks a Senior Site Reliability Engineer to boost reliability and scalability of compute provisioning and orchestration across data centers. You will work on Kubernetes, automation, observability, and incident response to lay a foundation for customer AI workloads.

Responsibilities include building monitoring and SLOs, automating remediation, and leading high-severity incidents.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior SRE: AI Cloud Platform & Kubernetes
Senior SRE: AI Cloud Platform & Kubernetes

Lambda Inc. • San Francisco (CA)

Hybrid
USD 190,000 - 270,000
Health insurance
401k with company match
Flexible PTO
+2
Senior SRE — Core Cloud Platform Resilience & Automation
Senior SRE — Core Cloud Platform Resilience & Automation

Front Door Defense • San Jose (CA)

On-site
USD 240,000 - 356,000
Health, dental, and vision coverage
401k Plan with company match
Flexible paid time off
+1
Senior SRE: Managed Kubernetes for AI Cloud
Senior SRE: Managed Kubernetes for AI Cloud

Lambda Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Health insurance
Dental insurance
Vision insurance
+3
Senior SRE: Managed Kubernetes for AI Cloud Platforms
Senior SRE: Managed Kubernetes for AI Cloud Platforms

Lambda • Bellevue (WA)

On-site
USD 240,000 - 356,000
Health, dental, and vision coverage
401k matching
Flexible PTO
+1
Senior SRE - Managed Kubernetes for AI Cloud
Senior SRE - Managed Kubernetes for AI Cloud

Socket.dev • San Francisco (CA)

Hybrid
USD 150,000 - 230,000
Health, dental, vision
4-day in-office work week
Wellness stipend
+1
Senior AI Cloud SRE — HPC & GPU Infra
Senior AI Cloud SRE — HPC & GPU Infra

Lambda Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Health insurance
Dental insurance
Vision insurance
+4
Senior SRE: Bare-Metal Kubernetes for AI Cloud
Senior SRE: Bare-Metal Kubernetes for AI Cloud

AI Chopping Block • San Francisco (CA), Northern (KY)

Hybrid
USD 190,000 - 270,000
Health, dental, vision
401k with company match
Wellness stipends
+1
Senior Kubernetes SRE — Scale AI Clusters & Automation
Senior Kubernetes SRE — Scale AI Clusters & Automation

Neura Market • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Equity compensation
Health, dental and vision coverage
Wellness and commuter stipends
+2
Senior AI Cloud SRE — Hybrid (SF/Bellevue)
Senior AI Cloud SRE — Hybrid (SF/Bellevue)

AI Chopping Block • San Francisco (CA), Northern (KY)

Hybrid
USD 170,000 - 250,000
Health, dental, vision coverage
Wellness stipend
Commuter stipend
+2
Senior Platform Engineer – Kubernetes & Cloud Infra
Senior Platform Engineer – Kubernetes & Cloud Infra

Lambda • San Francisco (CA)

Hybrid
USD 180,000 - 240,000
Health, dental, vision coverage
401k with company match
Wellness and commuter stipends
+1