Senior Storage Engineer

REALM

United States

Remote

USD 180,000 - 240,000

Full time

11 days ago
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

REALM is seeking a senior, hands-on storage architect to own the Ceph-based infrastructure powering petabyte-scale storage for GPU-intensive workloads. You will oversee architecture, deployment, and day-to-day operation across object, block, and file services.

You will mentor engineers, drive technical direction, and collaborate with compute/network teams, while participating in a shared on-call rotation with compensation.

Qualifications

  • Strong practical experience with Ceph covering architecture, deployment, troubleshooting, day-to-day operations, and optimisation.
  • Experience running distributed storage infrastructure at significant production scale.
  • A history of taking ownership of an infrastructure or storage domain and driving its technical direction.
  • Experience supporting and developing other engineers through mentoring and technical collaboration

Responsibilities

  • Own the architecture and production lifecycle of Ceph-based storage infrastructure at very large scale.
  • Help develop and operate an S3-compatible object storage offering, including access control, bucket and object workflows, reliability targets, observability and incident response
  • Work across the Ceph stack, including RADOS Gateway, RBD, and CephFS
  • Provide technical mentorship through design discussions, reviews, pairing, and participation in the on-call rotation
  • Evolve the storage platform as capacity grows from the petabyte range into substantially larger deployments
  • Partner with infrastructure teams responsible for compute and networking to ensure storage meets the needs of high-performance and GPU-heavy environments

Skills

Ceph architecture
Ceph deployment
Ceph troubleshooting
Distributed storage
Linux fundamentals
Networking fundamentals
On-call rotation
Mentoring engineers
Ownership of infrastructure

Tools

RADOS Gateway
RBD
CephFS

Job description

A fast-growing Neocloud - running the GPU infrastructure behind the next wave of AI companies, with storage as the foundation the whole platform sits on.

Strong revenue, well-funded, and growing fast - profitable and customer-backed well before the GPU buildout began.

The Role

This is a senior, hands-on role with broad technical ownership. You’ll take responsibility for the architecture, deployment, and day-to-day operation of large Ceph environments storing customer data at petabyte scale across object, block, and file services. The infrastructure supports demanding GPU and compute workloads, and you’ll play a key role in defining the direction of the storage platform and making the major technical decisions around it.

Specifically:
  • Own the architecture and production lifecycle of Ceph-based storage infrastructure at very large scale
  • Help develop and operate an S3-compatible object storage offering, including access control, bucket and object workflows, reliability targets, observability and incident response
  • Work across the Ceph stack, including RADOS Gateway, RBD, and CephFS
  • Provide technical mentorship through design discussions, reviews, pairing, and participation in the on-call rotation
  • Evolve the storage platform as capacity grows from the petabyte range into substantially larger deployments
  • Partner with infrastructure teams responsible for compute and networking to ensure storage meets the needs of high-performance and GPU-heavy environments
Who they’re looking for
  • Strong practical experience with Ceph covering architecture, deployment, troubleshooting, day-to-day operations, and optimisation
  • Experience running distributed storage infrastructure at significant production scale
  • A history of taking ownership of an infrastructure or storage domain and driving its technical direction
  • Experience supporting and developing other engineers through mentoring and technical collaboration
  • Solid Linux and networking fundamentals, with the ability to investigate complex issues across distributed systems
  • Willingness to participate in a shared production support/on-call rotation
  • The storage function is expanding, with a combination of experienced and developing engineers
  • The role is remote
  • Production support is shared across the team through a regular on-call schedule, with additional compensation for on-call participation
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Remote Senior Ceph Storage Engineer for Petabyte GPU Infra
Remote Senior Ceph Storage Engineer for Petabyte GPU Infra

REALM • United States

Remote
USD 180,000 - 240,000
Senior Storage Engineer
Senior Storage Engineer

Hydra Host • Miami (FL)

On-site
USD 120,000 - 160,000
Storage/System Engineers (Mid, Sr, Lead) – Ceph, Linux, Cloud & Enterprise Storage | 150+ PB scale | Mission-Critical Infrastructure at Global Scale FinTech
Storage/System Engineers (Mid, Sr, Lead) – Ceph, Linux, Cloud & Enterprise Storage | 150+ PB scale | Mission-Critical Infrastructure at Global Scale FinTech

Sprint and Partners • Illinois

On-site
USD 140,000 - 170,000
Storage/System Engineers (Mid, Sr, Lead) – Ceph, Linux, Cloud & Enterprise Storage | 150+ PB scale | Mission-Critical Infrastructure at Global Scale FinTech
Storage/System Engineers (Mid, Sr, Lead) – Ceph, Linux, Cloud & Enterprise Storage | 150+ PB scale | Mission-Critical Infrastructure at Global Scale FinTech

Sprint and Partners • San Francisco (CA)

On-site
USD 140,000 - 190,000
Senior Storage Engineer
Senior Storage Engineer

Kindredventures • Miami (FL)

On-site
USD 120,000 - 150,000
Senior Staff Storage Engineer
Senior Staff Storage Engineer

DDN • Sacramento (CA)

On-site
USD 180,000 - 240,000
Cloud Engineer
Cloud Engineer

Cirrascale Corporation • San Diego (CA)

Hybrid
USD 121,000 - 179,000
Comprehensive medical, dental, vision coverage
401(k) with company match
Generous paid time-off
+1
Senior Storage Software Engineer - DGX Cloud
Senior Storage Software Engineer - DGX Cloud

NVIDIA • Santa Clara (CA)

On-site
USD 224,000 - 431,000
Equity
Benefits
Senior Storage Platform Engineer
Senior Storage Platform Engineer

Nvidia Corporation in • Santa Clara (CA)

Hybrid
USD 208,000 - 334,000
Equity
Benefits
Senior Manager, Storage Engineering
Senior Manager, Storage Engineering

NVIDIA • Santa Clara (CA)

On-site
USD 248,000 - 397,000
Equity
Benefits