Senior Storage Reliability Engineer – Hybrid (SF/SJ)

Neura Market

San Francisco, Northern (CA, KY)

Hybrid

USD 180,000 - 240,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Health, dental, vision coverage
401(k) with company match
Flexible paid time off
Wellness and commuter stipends

Job summary

Lambda is seeking a Storage Engineer to own reliability, performance, and capacity for the production storage fleet across data centers. You will build monitoring dashboards, investigate incidents, and automate workflows to reduce toil while enabling self-healing capabilities.

The role requires hands-on SD storage experience at scale, strong incident-response skills, and collaboration with cross-functional teams to deploy storage solutions across sites.

Qualifications

  • 5+ years of experience operating Linux systems in production or HPC environments with scale-out or software-defined storage.

Responsibilities

  • Own the reliability, performance, and capacity health of Lambda's production storage fleet across data centers.
  • Build and maintain monitoring, dashboards, and alerting for storage performance, capacity, and hardware failures.
  • Investigate and resolve storage incidents using telemetry, logs, and performance profiling.
  • Automate ticketing, escalation, and incident-response workflows to reduce triage time.
  • Design and maintain self-healing automation for drive replacement, node swaps, and capacity rebalancing.
  • Implement CI/CD pipelines for storage automation and tooling.
  • Collaborate with Storage Engineers and Release teams to deploy software-defined storage across sites using Ansible, Jenkins, etc.
  • Work with hardware and networking teams to diagnose low-level I/O and network issues affecting storage.
  • Participate in an on-call rotation to reduce MTTR and improve automation.

Skills

Linux systems
Storage at scale
Monitoring & logging
Kubernetes
CI/CD pipelines
Infrastructure as Code
Ansible
Python or Go
Networking basics
Incident response

Tools

Prometheus
Grafana
Alertmanager
Datadog
SumoLogic
ArgoCD
Helm
Kustomize
GitHub Actions
Jenkins
BuildKite
Docker
Podman
Terraform
Ansible

Job description

Lambda is seeking a Storage Engineer to own reliability, performance, and capacity for the production storage fleet across data centers. You will build monitoring dashboards, investigate incidents, and automate workflows to reduce toil while enabling self-healing capabilities.

The role requires hands-on SD storage experience at scale, strong incident-response skills, and collaboration with cross-functional teams to deploy storage solutions across sites.

Get your free, confidential resume review.
or drag and drop your file here.