Senior Storage Reliability Engineer - AI Cloud Infra

Lambda Inc.

San Francisco, Northern (CA, KY)

Hybrid

USD 180,000 - 240,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Wellness stipend
Commuter stipend
401k match

Job summary

Lambda Inc. in San Francisco is seeking a Storage Engineer to own the reliability, performance, and capacity of our production storage fleet across multiple data centers, using a software-defined data plane.

You will build monitoring, dashboards, and alerting for storage performance and hardware failures, and automate workflows to speed root-cause analysis. A strong background in SDS, Kubernetes, CI/CD, and IaC is essential, with on-call responsibilities and deep telemetry-driven incident

Qualifications

  • 5+ years of Linux production/HPC experience with scale-out storage on software-defined platforms.
  • Hands-on experience with Software-Defined Storage (SDS) platforms at scale.
  • Strong incident-response instincts from alert to postmortem.
  • Experience with monitoring/logging platforms (Prometheus, Grafana, Alertmanager, Datadog, SumoLogic).
  • Kubernetes experience with GitOps tooling (ArgoCD, Helm/Kustomize).
  • CI/CD tooling (GitHub Actions, Jenkins, BuildKite), containerization (Docker/Podman), and Python or Go.
  • Infrastructure as Code (Terraform, Ansible).
  • Storage protocols across file (NFS/SMB), object (S3), block (NVMe-oF/TCP), and structured storage (vector DB, SQL).

Responsibilities

  • Own reliability, performance, and capacity health of Lambda's production storage fleet across data centers.
  • Build and maintain monitoring, dashboards, and alerting for storage performance, capacity, and hardware failures.
  • Investigate and resolve storage incidents using telemetry, logs, and performance profiling, from NIC issues to cluster rebuilds.
  • Automate ticketing, escalation, and incident-response workflows to reduce redundant triage.
  • Design and maintain self-healing automation for drive replacement, node swaps, rebuild monitoring, and capacity rebalancing.
  • Implement CI/CD pipelines for storage automation and tooling.
  • Collaborate with Storage Engineers, Fleet Orchestration, and Release Engineering to deploy and configure software-defined storage across sites using Ansible, Jenkins, etc.
  • Work with hardware and networking teams to diagnose low-level I/O and network issues surface as storage symptoms.
  • Participate in on-call rotation to reduce MTTR and build automation to avoid repeat paging.

Skills

Linux admin
Scale-out storage
SDS platforms
Monitoring dashboards
Kubernetes
CI/CD tooling
IaC tools
Networking troubleshooting
Python/Go
Jenkins/GitHub Actions

Tools

Prometheus
Grafana
Alertmanager
Datadog
SumoLogic
ArgoCD
Helm/Kustomize
Docker/Podman
Terraform
Ansible

Job description

Lambda Inc. in San Francisco is seeking a Storage Engineer to own the reliability, performance, and capacity of our production storage fleet across multiple data centers, using a software-defined data plane.

You will build monitoring, dashboards, and alerting for storage performance and hardware failures, and automate workflows to speed root-cause analysis. A strong background in SDS, Kubernetes, CI/CD, and IaC is essential, with on-call responsibilities and deep telemetry-driven incident

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Storage Reliability Engineer, Hybrid (SF/SJ)
Senior Storage Reliability Engineer, Hybrid (SF/SJ)

Lambda • San Jose (CA)

On-site
USD 150,000 - 190,000
Health, dental, and vision coverage
401k with 2% company match
Flexible paid time off
+1
Senior Storage Reliability Engineer — AI Cloud (Onsite 4x/wk)
Senior Storage Reliability Engineer — AI Cloud (Onsite 4x/wk)

Socket.dev • San Francisco (CA)

On-site
USD 140,000 - 210,000
Health coverage
Dental coverage
Vision coverage
+3
Senior Storage SRE — Scalable AI Cloud
Senior Storage SRE — Scalable AI Cloud

Artha Nexgen • Chicago (IL), Northern (KY)

Hybrid
USD 267,000 - 356,000
Health, dental, vision coverage
Wellness stipend
401k with company match
+1
Senior Storage Reliability Engineer – Scale-Out AI Cloud
Senior Storage Reliability Engineer – Scale-Out AI Cloud

Lambda • San Francisco (CA)

Hybrid
USD 267,000 - 356,000
Health, dental, and vision
Wellness stipend
401k with company match
+1
Senior Storage Reliability Engineer – Hybrid (SF/SJ)
Senior Storage Reliability Engineer – Hybrid (SF/SJ)

Neura Market • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Health, dental, vision coverage
401(k) with company match
Flexible paid time off
+1
Senior Storage Systems Engineer - AI Cloud Infra
Senior Storage Systems Engineer - AI Cloud Infra

Lambda Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Health insurance
Dental insurance
Vision insurance
+3
Senior Storage SRE - Scale, Automation & Reliability
Senior Storage SRE - Scale, Automation & Reliability

Lambda • United States

Remote
USD 180,000 - 240,000
Senior Storage Engineer - AI Cloud Infrastructure
Senior Storage Engineer - AI Cloud Infrastructure

Lambda • San Francisco (CA)

On-site
USD 180,000 - 260,000
Cash & equity
Health, dental, vision
Wellness stipend
+2
Staff Storage Engineer — Scale AI Infrastructure
Staff Storage Engineer — Scale AI Infrastructure

Lambda • San Francisco (CA)

On-site
USD 314,000 - 465,000
Health, dental, and vision coverage
401k Plan with 2% company match
Flexible paid time off plan
Senior Software Engineer – AI Storage & Kubernetes
Senior Software Engineer – AI Storage & Kubernetes

Lambda • United States

Hybrid
USD 160,000 - 210,000