Senior Storage Reliability Engineer — AI Cloud (Onsite 4x/wk)

Socket.dev

San Francisco (CA)

On-site

USD 140,000 - 210,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Health coverage
Dental coverage
Vision coverage
401k with match
Wellness stipend
Commuter stipend

Job summary

Lambda is building the world's best AI cloud. We seek a Storage Engineer to own reliability and capacity for our production storage fleet, spanning on-site data centers and a software-defined data plane. You will automate, monitor, and optimize performance across multiple sites, collaborating with hardware, networking, and release teams.

You will drive incident response, build dashboards, and implement self-healing automation while advancing CI/CD for storage tooling and integrations.

Qualifications

  • 5+ years operating Linux systems in production or HPC environments.
  • Hands-on storage experience at scale on software-defined platforms.
  • Strong incident-response ownership from alert to postmortem.
  • Experience with monitoring/logging: Prometheus, Grafana, Datadog, Sumo Logic.
  • Kubernetes experience including GitOps tools like ArgoCD, Helm, Kustomize.
  • CI/CD tooling: GitHub Actions, Jenkins, BuildKite; Docker/Podman.
  • Programming in Python or Go; Infrastructure as Code with Terraform/Ansible.
  • Core storage protocols: NFS, SMB, S3, NVMe-oF; basic DB concept.

Responsibilities

  • Own reliability, performance, and capacity health of production storage fleet.
  • Build and maintain monitoring, dashboards, and alerting for storage systems.
  • Investigate and resolve storage incidents using telemetry and profiling.
  • Automate ticketing, escalation, and incident-response workflows.
  • Design self-healing automation for drive replacements, node swaps, rebuilds.
  • Implement CI/CD pipelines for storage automation and tooling.
  • Collaborate with Storage Engineers, Fleet Orchestration, Release Engineering.
  • Work with hardware and networking teams to diagnose low-level I/O issues.
  • Participate in on-call rotations to minimize MTTR.

Skills

Linux production systems
Storage at scale (SDS)
Incident response
Monitoring dashboards
Kubernetes
CI/CD tooling
Python/Go
Infrastructure as code
Storage protocols

Education

Bachelor's degree in Computer Science or related field

Tools

Prometheus
Grafana
Alertmanager
Datadog
Sumo Logic
Kubernetes
ArgoCD
Helm
Kustomize
GitHub Actions
Jenkins
BuildKite
Docker
Terraform
Ansible

Job description

Lambda is building the world's best AI cloud. We seek a Storage Engineer to own reliability and capacity for our production storage fleet, spanning on-site data centers and a software-defined data plane. You will automate, monitor, and optimize performance across multiple sites, collaborating with hardware, networking, and release teams.

You will drive incident response, build dashboards, and implement self-healing automation while advancing CI/CD for storage tooling and integrations.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Storage Reliability Engineer - AI Cloud Infra
Senior Storage Reliability Engineer - AI Cloud Infra

Lambda Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Wellness stipend
Commuter stipend
401k match
Senior Storage Reliability Engineer – Scale-Out AI Cloud
Senior Storage Reliability Engineer – Scale-Out AI Cloud

Lambda • San Francisco (CA)

Hybrid
USD 267,000 - 356,000
Health, dental, and vision
Wellness stipend
401k with company match
+1
Senior Storage SRE — Scalable AI Cloud
Senior Storage SRE — Scalable AI Cloud

Artha Nexgen • Chicago (IL), Northern (KY)

Hybrid
USD 267,000 - 356,000
Health, dental, vision coverage
Wellness stipend
401k with company match
+1
Senior Storage Engineer - AI Cloud Infrastructure
Senior Storage Engineer - AI Cloud Infrastructure

Lambda • San Francisco (CA)

On-site
USD 180,000 - 260,000
Cash & equity
Health, dental, vision
Wellness stipend
+2
Senior Storage SRE - Scale, Automation & Reliability
Senior Storage SRE - Scale, Automation & Reliability

Lambda • United States

Remote
USD 180,000 - 240,000
Senior Storage Reliability Engineer, Hybrid (SF/SJ)
Senior Storage Reliability Engineer, Hybrid (SF/SJ)

Lambda • San Jose (CA)

On-site
USD 150,000 - 190,000
Health, dental, and vision coverage
401k with 2% company match
Flexible paid time off
+1
Senior Storage Systems Engineer - AI Cloud Infra
Senior Storage Systems Engineer - AI Cloud Infra

Lambda Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Health insurance
Dental insurance
Vision insurance
+3
Senior Storage Reliability Engineer – Hybrid (SF/SJ)
Senior Storage Reliability Engineer – Hybrid (SF/SJ)

Neura Market • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Health, dental, vision coverage
401(k) with company match
Flexible paid time off
+1
Staff Storage Engineer — Scale AI Infrastructure
Staff Storage Engineer — Scale AI Infrastructure

Lambda • San Francisco (CA)

On-site
USD 314,000 - 465,000
Health, dental, and vision coverage
401k Plan with 2% company match
Flexible paid time off plan
Senior Site Reliability Engineer - AI Cloud Platform
Senior Site Reliability Engineer - AI Cloud Platform

Lambda • United States

Hybrid
USD 160,000 - 220,000