Remote Senior Storage Engineer - AI Infrastructure

Runpod

Mount Laurel Township (NJ)

On-site

USD 180,000 - 260,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Equity participation
Medical, dental & vision plans
Flexible PTO
Remote-friendly / Slack-native culture
Home office stipend

Job summary

Runpod is seeking a Senior Storage Engineer to design, scale, and operate its high-performance storage platform. You’ll own capacity, durability, and performance of network volumes, local NVMe, and S3-compatible object storage, writing production code and automating operations for petabyte-scale workloads.

Join a remote-first, Slack-native team that emphasizes rapid delivery, ownership, and collaboration with SRE, networking, and hardware partners to reduce cold starts and optimize data movement

Qualifications

  • 8+ years in infrastructure, storage, or systems engineering with production storage ownership.
  • Deep, practical experience with at least one distributed storage system (Ceph, MinIO, Lustre, GPFS/Spectrum Scale, MooseFS, WekaFS, VAST, or similar).
  • Strong Linux internals and storage-stack knowledge: block layer, filesystems, NVMe, I/O schedulers, NFS/SMB, iSCSI/NVMe-oF.
  • Experience building and operating S3-compatible object storage services.
  • Solid networking fundamentals tuned for storage workloads.
  • Production-code ability in Go, Python, Rust, or similar (not just scripting).
  • Hands-on observability tooling (Prometheus, Grafana, Datadog) with designing metrics.
  • Ownership, self-starting, collaborative, low-ego attitude.

Responsibilities

  • Own capacity, durability, availability, and performance characteristics of storage tiers.
  • Tune the full I/O path: device and filesystem configuration, caching and read-ahead strategies, replication and erasure coding trade-offs, and client-side mount behavior.
  • Diagnose hard performance problems end to end.
  • Lead capacity expansions, migrations, and rebalances with minimal customer-visible disruption.
  • Work with Runpod and partner networking teams to design and tune storage paths for high throughput.
  • Understand and optimize RDMA/RoCE and high-speed fabrics for storage traffic.
  • Collaborate with network engineering on topology decisions and cross-region data movement.
  • Write production code (Go, Python, or similar) for storage control-plane services and data pipelines.
  • Build against and extend APIs: control plane, S3 interfaces, CSI, Kubernetes, vendor/cloud APIs.
  • Automate operations; treat infrastructure as code and participate in CI.
  • Instrument storage fleet with metrics and dashboards; establish SLOs and alerts.
  • Participate in on-call rotation and blameless post-incident reviews.

Skills

Distributed storage
Linux internals
Networking fundamentals
Go
Python
Observability

Tools

Ceph
MinIO
Lustre
GPFS/Spectrum Scale
MooseFS
WekaFS
VAST
ZFS-based systems

Job description

Runpod is seeking a Senior Storage Engineer to design, scale, and operate its high-performance storage platform. You’ll own capacity, durability, and performance of network volumes, local NVMe, and S3-compatible object storage, writing production code and automating operations for petabyte-scale workloads.

Join a remote-first, Slack-native team that emphasizes rapid delivery, ownership, and collaboration with SRE, networking, and hardware partners to reduce cold starts and optimize data movement

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

HPC Storage Engineer - West Coast
HPC Storage Engineer - West Coast

Runpod • Mount Laurel Township (NJ)

On-site
USD 180,000 - 260,000
Equity participation
Medical, dental & vision plans
Flexible PTO
+2
Head of Scalable AI Infrastructure & SRE
Head of Scalable AI Infrastructure & SRE

Runpod • United States

On-site
USD 225,000 - 325,000
Stock options
Remote-first culture
Competitive salary
+2
Remote SRE: AI Platform Reliability & Automation
Remote SRE: AI Platform Reliability & Automation

Runpod • United States

On-site
USD 150,000 - 200,000
Remote work first
Competitive base salary
Stock options equity
+2
Global AI Infra Architect & SRE Leader
Global AI Infra Architect & SRE Leader

Runpod • Mount Laurel Township (NJ)

On-site
USD 225,000 - 325,000
Equity
Medical/Dental/Vision
Flexible PTO
+3
Senior Data Engineer - Build Scalable Data Platform (Remote)
Senior Data Engineer - Build Scalable Data Platform (Remote)

Runpod • United States

Remote
USD 175,000 - 220,000
Equity
Flexible PTO
Remote-first
+1
Engineering Manager, Cloud AI Platform
Engineering Manager, Cloud AI Platform

RunPod Inc. • United States

On-site
USD 110,000 - 220,000
Equity
Medical benefits
Flexible PTO
+2
Senior Product Manager, Remote AI Infrastructure
Senior Product Manager, Remote AI Infrastructure

Runpod • Mount Laurel Township (NJ)

On-site
USD 175,000 - 225,000
Meaningful equity
Medical, dental & vision plans
Flexible PTO
+2
Senior Data Engineer - AI-Driven Data Pipelines (Remote-First)
Senior Data Engineer - AI-Driven Data Pipelines (Remote-First)

Runpod • Mount Laurel Township (NJ)

On-site
USD 175,000 - 220,000
Equity
Remote-friendly
Flexible PTO
+2
Remote Full-Stack Software Engineer for AI Infrastructure
Remote Full-Stack Software Engineer for AI Infrastructure

Grapevine Round1 AI • Northern (KY)

Hybrid
USD 110,000 - 170,000
Senior Data Engineer - Remote, Equity & Flexible PTO
Senior Data Engineer - Remote, Equity & Flexible PTO

Runpod • San Francisco (CA)

On-site
USD 175,000 - 220,000
Equity
Flexible PTO
Remote-first culture
+1