Senior SRE: Managed Kubernetes for AI Cloud Platforms

Lambda

Bellevue (WA)

On-site

USD 240,000 - 356,000

Full time

13 days ago
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Health, dental, and vision coverage
401k matching
Flexible PTO
Wellness and commuter stipends

Job summary

Lambda, The Superintelligence Cloud, seeks an experienced SRE to operate and scale Kubernetes clusters in production, collaborating with HPC Ops and Datacenter Ops. This role requires presence in Bellevue/SF/SJ offices four days per week, with a designated WFH day on Tuesday.

You will design control planes, automate lifecycle tasks, and define SLOs/SLIs while supporting customers during incidents. Proficiency in Go, Python, and Kubernetes tooling is essential.

Qualifications

  • 6+ years in SRE/ops or similar role with Linux cluster expertise.
  • Strong Go and Python skills; experience with GitOps (ArgoCD) and Helm.
  • Production Kubernetes experience (on-prem, EKS, GKE, or similar).
  • Able to work independently or in a team; strong customer incident handling and communication.
  • Familiarity with observability tools (Prometheus, Grafana, FluentBit) and CI/CD pipelines.

Responsibilities

  • Operate and scale bare-metal Kubernetes clusters up to thousands of nodes.
  • Handle cluster degradation, recovery, resizing, and incident response.
  • Participate in a well-managed on-call rotation for critical incidents.
  • Assist customers with Kubernetes questions, workload integration, storage, and authentication.
  • Collaborate with HPC Ops and Datacenter Ops for cross-functional issues.
  • Use Python and Go to create tooling and automate quality validation.
  • Design and maintain scalable control plane services, operators, and controllers.
  • Develop automation for cluster lifecycle management: provisioning, upgrades, patching, deletion.
  • Define and implement SLOs/SLIs for Kubernetes services and platform reliability.

Skills

Go
Python
GitOps
Kubernetes
Incident response
Observability

Tools

ArgoCD
Helm
Kubernetes operators
Prometheus
Grafana
Fluent Bit
Cluster API
kubeadm

Job description

Lambda, The Superintelligence Cloud, seeks an experienced SRE to operate and scale Kubernetes clusters in production, collaborating with HPC Ops and Datacenter Ops. This role requires presence in Bellevue/SF/SJ offices four days per week, with a designated WFH day on Tuesday.

You will design control planes, automate lifecycle tasks, and define SLOs/SLIs while supporting customers during incidents. Proficiency in Go, Python, and Kubernetes tooling is essential.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior SRE: Managed Kubernetes for AI Cloud (Hybrid)
Senior SRE: Managed Kubernetes for AI Cloud (Hybrid)

Lambda • San Francisco (CA)

Hybrid
USD 170,000 - 260,000
401k Plan with company match (USA)
Health, dental, and vision coverage
Wellness and commuter stipends
+1
Senior SRE: Managed Kubernetes for AI Cloud
Senior SRE: Managed Kubernetes for AI Cloud

Lambda Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Health insurance
Dental insurance
Vision insurance
+3
Senior SRE - Managed Kubernetes for AI Cloud
Senior SRE - Managed Kubernetes for AI Cloud

Socket.dev • San Francisco (CA)

Hybrid
USD 150,000 - 230,000
Health, dental, vision
4-day in-office work week
Wellness stipend
+1
Senior SRE: AI Cloud Platform & Kubernetes
Senior SRE: AI Cloud Platform & Kubernetes

Lambda Inc. • San Francisco (CA)

Hybrid
USD 190,000 - 270,000
Health insurance
401k with company match
Flexible PTO
+2
Senior SRE: AI Cloud Platform & Kubernetes Expert
Senior SRE: AI Cloud Platform & Kubernetes Expert

Lambda • Bellevue (WA)

On-site
USD 180,000 - 260,000
Health insurance
Dental insurance
Vision insurance
+3
Senior SRE – AI Cloud Platform, Kubernetes Expert
Senior SRE – AI Cloud Platform, Kubernetes Expert

Socket.dev • San Francisco (CA)

On-site
USD 180,000 - 240,000
Health, dental, vision coverage for in
Wellness and commuter stipends
401k with 2% company match
+1
Senior SRE: Bare-Metal Kubernetes for AI Cloud
Senior SRE: Bare-Metal Kubernetes for AI Cloud

AI Chopping Block • San Francisco (CA), Northern (KY)

Hybrid
USD 190,000 - 270,000
Health, dental, vision
401k with company match
Wellness stipends
+1
Senior Kubernetes SRE — Scale AI Clusters & Automation
Senior Kubernetes SRE — Scale AI Clusters & Automation

Neura Market • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Equity compensation
Health, dental and vision coverage
Wellness and commuter stipends
+2
Senior AI Cloud SRE — HPC & GPU Reliability Lead
Senior AI Cloud SRE — HPC & GPU Reliability Lead

Lambda • San Francisco (CA)

On-site
USD 180,000 - 230,000
Health, dental, and vision
401k with company match
Flexible paid time off
+2
Senior AI Cloud SRE — Hybrid (SF/Bellevue)
Senior AI Cloud SRE — Hybrid (SF/Bellevue)

AI Chopping Block • San Francisco (CA), Northern (KY)

Hybrid
USD 170,000 - 250,000
Health, dental, vision coverage
Wellness stipend
Commuter stipend
+2