Senior Storage Reliability Engineer for AI Cloud

crusoe

Sunnyvale (CA)

On-site

USD 170,000 - 205,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

RSUs
Health insurance
HSA
Parental leave
Disability insurance
401(k) match
PTO
Cell phone reimbursement
Tuition reimbursement

Job summary

Crusoe Energy Systems in Sunnyvale, CA, is seeking a Storage SRE to maintain and scale our AI-ready cloud storage. You will build automation and self-healing tools to monitor Crusoe's distributed storage infrastructure, including block, file, and object storage.

You will drive reliability initiatives for data replication, encryption, backup, and failover. Work with storage engineers, kernel and hardware teams to optimize I/O paths, cache policies, and NVMe-/SSD-backed volumes across large AI

Qualifications

  • Bachelor's degree in Computer Science, Electrical Engineering, or related field; equivalent practical experience
  • 5+ years of Storage SRE, systems, or storage engineering experience
  • Hands-on with enterprise storage platforms (Pure Storage, EMC) and understanding of storage architectures
  • Deep knowledge of object, block, and file storage paradigms
  • Proficiency in Go, Python, Java, or C
  • Experience with Infrastructure as Code and deployment tools (Terraform, Ansible, Puppet)
  • Strong Linux internals knowledge focusing on I/O, memory, and storage scheduling
  • Familiarity with NFS, SMB, iSCSI, NVMe-oF
  • Experience with containerized workloads and orchestration (Kubernetes, Docker)
  • Excellent incident response, troubleshooting, and documentation practices
  • Experience with managed storage services on AWS, GCP, Azure

Responsibilities

  • Build automation and self-healing tools to monitor Crusoe's distributed cloud storage infrastructure
  • Drive reliability initiatives for data replication, encryption, backup, and robust failover mechanisms
  • Collaborate with storage engineers and hardware/kernel teams to optimize I/O paths and storage scheduling
  • Support high-performance NVMe- and SSD-backed volumes for large-scale AI compute clusters
  • Contribute to architecture of fault-tolerant, scalable storage backends for AI-first cloud environments

Skills

Go
Python
Java
C
Terraform
Ansible
Puppet
Linux internals
Storage protocols
Kubernetes
Docker
Incident response

Education

Bachelor's degree in CS/EE or related

Tools

Terraform
Ansible
Puppet

Job description

Crusoe Energy Systems in Sunnyvale, CA, is seeking a Storage SRE to maintain and scale our AI-ready cloud storage. You will build automation and self-healing tools to monitor Crusoe's distributed storage infrastructure, including block, file, and object storage.

You will drive reliability initiatives for data replication, encryption, backup, and failover. Work with storage engineers, kernel and hardware teams to optimize I/O paths, cache policies, and NVMe-/SSD-backed volumes across large AI

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Storage SRE - Scalable AI Cloud & Data Platform
Senior Storage SRE - Scalable AI Cloud & Data Platform

Crusoe Energy Systems LLC • Sunnyvale (CA)

On-site
USD 170,000 - 205,000
Industry competitive pay
Restricted Stock Units
Health insurance options (HDHP/PPO)
+9
Senior AI Cloud Platform Engineer - Reliability & Scale
Senior AI Cloud Platform Engineer - Reliability & Scale

Crusoe • San Francisco (CA)

On-site
USD 170,000 - 205,000
Health insurance
RSUs
401(k) match
+2
Senior Production Engineer, Storage
Senior Production Engineer, Storage

crusoe • Sunnyvale (CA)

On-site
USD 170,000 - 205,000
RSUs
Health insurance
HSA
+6
Senior Production Engineer, Storage
Senior Production Engineer, Storage

Crusoe Energy Systems LLC • Sunnyvale (CA)

On-site
USD 170,000 - 205,000
Industry competitive pay
Restricted Stock Units
Health insurance options (HDHP/PPO)
+9
Cloud Storage Engineer I: Grow in Distributed Systems
Cloud Storage Engineer I: Grow in Distributed Systems

Crusoe Energy Systems • San Francisco (CA)

On-site
USD 100,000 - 140,000
Health & wellbeing benefits
Paid time off
401(k) match
+1
Senior Storage SRE — Scalable AI Cloud
Senior Storage SRE — Scalable AI Cloud

Artha Nexgen • Chicago (IL), Northern (KY)

Hybrid
USD 267,000 - 356,000
Health, dental, vision coverage
Wellness stipend
401k with company match
+1
Senior Compute Engineer: AI Infra, Kernel & Virtualization
Senior Compute Engineer: AI Infra, Kernel & Virtualization

crusoe • Sunnyvale (CA)

On-site
USD 170,000 - 205,000
Health insurance
RSUs
401(k) match
+7
Senior Storage Reliability Engineer - AI Cloud Infra
Senior Storage Reliability Engineer - AI Cloud Infra

Lambda Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Wellness stipend
Commuter stipend
401k match
Staff Compute Systems Engineer, AI Infrastructure
Staff Compute Systems Engineer, AI Infrastructure

Crusoe • San Francisco (CA)

On-site
USD 209,000 - 253,000
RSUs
PTO & holidays
Health insurance
+2
Senior Storage SRE - Scale, Automation & Reliability
Senior Storage SRE - Scale, Automation & Reliability

Lambda • United States

Remote
USD 180,000 - 240,000