Senior Storage Reliability Engineer – Scale-Out AI Cloud

Lambda

San Francisco (CA)

Hybrid

USD 267,000 - 356,000

Full time

12 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Health, dental, and vision
Wellness stipend
401k with company match
Flexible PTO

Job summary

Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure powering AI researchers to hyperscalers. Our Storage Engineering team builds the data platform and software-defined storage powering Lambda’s world-class services.

This role focuses on reliability, performance, capacity, monitoring, and automation. You’ll own incidents end to end, design self-healing systems, and collaborate with cross-functional teams to deploy storage at scale across data centers.

Qualifications

  • 5+ years of experience operating Linux systems in production or HPC environments with hands-on storage experience at scale.
  • Hands-on experience with Software-Defined Storage (SDS) platforms and integration with management and data-plane APIs.
  • Strong incident-response instincts, owning storage incidents end to end from alert to postmortem.
  • Experience with monitoring/logging platforms (Prometheus, Grafana, Alertmanager, Datadog, SumoLogic) and building dashboards/alerts.
  • Kubernetes experience with GitOps tooling (ArgoCD, Helm/Kustomize) and troubleshooting.
  • CI/CD tooling (GitHub Actions, Jenkins, BuildKite), containerization (Docker/Podman), and systems programming (Python or Go).
  • Infrastructure as Code (Terraform, Ansible).
  • Storage protocols across file, object, block, and structured storage.

Responsibilities

  • Own the reliability, performance, and capacity health of Lambda's production storage fleet across all data centers.
  • Build and maintain monitoring, dashboards, and alerting for storage performance, capacity, and hardware failures.
  • Investigate and resolve storage-related incidents using telemetry, logs, and profiling.
  • Automate ticketing, escalation, and incident-response workflows to reduce repetitive triage.
  • Design self-healing automation for drive replacements, node swaps, and capacity rebalancing.
  • Implement CI/CD pipelines for storage automation and tooling.
  • Collaborate with Storage Engineers, Fleet Orchestration, and Release Engineering to deploy software-defined storage across sites.
  • Work with hardware and networking teams to diagnose low-level I/O and network issues surface as storage symptoms.
  • Participate in on-call rotation to drive down MTTR and reduce repeat paging.

Skills

Linux systems
Storage at scale
SDS platforms
Monitoring dashboards
Kubernetes
CI/CD tooling
IaC Terraform/Ansible
Networking I/O

Tools

Prometheus
Grafana
GitHub Actions
Jenkins
BuildKite
Docker/Podman
Ansible
Terraform

Job description

Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure powering AI researchers to hyperscalers. Our Storage Engineering team builds the data platform and software-defined storage powering Lambda’s world-class services.

This role focuses on reliability, performance, capacity, monitoring, and automation. You’ll own incidents end to end, design self-healing systems, and collaborate with cross-functional teams to deploy storage at scale across data centers.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Storage Reliability Engineer - AI Cloud Infra
Senior Storage Reliability Engineer - AI Cloud Infra

Lambda Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Wellness stipend
Commuter stipend
401k match
Senior Storage Reliability Engineer — AI Cloud (Onsite 4x/wk)
Senior Storage Reliability Engineer — AI Cloud (Onsite 4x/wk)

Socket.dev • San Francisco (CA)

On-site
USD 140,000 - 210,000
Health coverage
Dental coverage
Vision coverage
+3
Senior Storage Engineer - AI Cloud Infrastructure
Senior Storage Engineer - AI Cloud Infrastructure

Lambda • San Francisco (CA)

On-site
USD 180,000 - 260,000
Cash & equity
Health, dental, vision
Wellness stipend
+2
Senior Storage SRE - Scale, Automation & Reliability
Senior Storage SRE - Scale, Automation & Reliability

Lambda • United States

Remote
USD 180,000 - 240,000
Senior Storage Reliability Engineer, Hybrid (SF/SJ)
Senior Storage Reliability Engineer, Hybrid (SF/SJ)

Lambda • San Jose (CA)

On-site
USD 150,000 - 190,000
Health, dental, and vision coverage
401k with 2% company match
Flexible paid time off
+1
Senior Software Engineer – AI Storage & Kubernetes
Senior Software Engineer – AI Storage & Kubernetes

Lambda • United States

Hybrid
USD 160,000 - 210,000
Senior Storage Systems Engineer - AI Cloud Infra
Senior Storage Systems Engineer - AI Cloud Infra

Lambda Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Health insurance
Dental insurance
Vision insurance
+3
Senior Software Engineer (Infrastructure Storage)
Senior Software Engineer (Infrastructure Storage)

Lambda • United States

Hybrid
USD 150,000 - 230,000
Staff Storage Engineer — Scale AI Infrastructure
Staff Storage Engineer — Scale AI Infrastructure

Lambda • San Francisco (CA)

On-site
USD 314,000 - 465,000
Health, dental, and vision coverage
401k Plan with 2% company match
Flexible paid time off plan
Senior Site Reliability Engineer - AI Cloud Platform
Senior Site Reliability Engineer - AI Cloud Platform

Lambda • United States

Hybrid
USD 160,000 - 220,000