Storage/System Engineers (Mid, Sr, Lead) – Ceph, Linux, Cloud & Enterprise Storage | 150+ PB scale | Mission-Critical Infrastructure at Global Scale FinTech

Sprint and Partners

San Francisco (CA)

On-site

USD 140,000 - 190,000

Full time

6 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Sprint and Partners seeks an experienced storage engineer to design, operate and optimize a Ceph-driven storage platform for a global fintech service. You will manage data placement, replication, and recovery while coordinating with database, networking and platform teams in a primarily on-premises environment.

The role emphasizes deep production experience with Ceph or enterprise storage, Linux fundamentals, and hands-on capabilities with storage hardware across multiple vendors.

Qualifications

  • Deep expertise in production Ceph or enterprise storage with exposure to the other.
  • Proven ability to troubleshoot active production clusters, not just deployments.
  • Strong Linux and distributed-storage fundamentals with willingness to deepen Ceph expertise.

Responsibilities

  • Design, operate and evolve storage infrastructure for a global fintech platform.
  • Reason about data placement, replication, quorum, failure domains, latency and capacity.
  • Trace complex problems across disks, controllers, networks and the Linux storage stack.
  • Collaborate with database, networking, application and platform engineering teams.
  • Operate SAN, NAS, DAS and object storage across Ceph and enterprise appliances.
  • Automate provisioning, configuration and upgrades using Bash, Python, Ansible or Puppet.

Skills

Ceph
Linux
Distributed storage
Automation
Python
Bash
Ansible
Puppet
Prometheus
Grafana
Networking
NVMe

Tools

Pure Storage
NetApp
Dell
Hitachi
HPE
IBM

Job description

This global fintech operates business-critical infrastructure at significant scale (200 PB+). Its storage platform supports workloads ranging from high-throughput, low-latency transactional databases to petabyte-scale object stores, making data integrity and availability fundamental to the product.

This is not a conventional public-cloud role. The core platform runs on on-premises, bare-metal infrastructure, and engineers own the full storage stack: hardware, Linux, distributed storage, automation, observability, capacity and failure recovery. The company is moving to Cloud which will be an exciting process as well.

The environment combines Ceph with enterprise platforms such as Pure Storage, NetApp, Dell, Hitachi and HPE. You may be strongest in either software-defined or enterprise storage, provided you have strong fundamentals and can develop further in the other domain.

Engineers currently focused on cloud infrastructure are also relevant if they have previous hands-on bare-metal or on-premises experience.

The role

  • Design, operate and evolve storage infrastructure supporting a global financial technology platform.
  • Reason about data placement, replication, quorum, failure domains, recovery behaviour, latency, throughput and capacity.
  • Trace complex problems across disks, controllers, networks, the Linux storage stack and Ceph.
  • Become the first Chicago-based engineer in an established Storage team of approximately six engineers, currently centred in Amsterdam.
  • Collaborate with database, networking, application and Platform Engineering teams.

What you’ll work on

  • Design and operate distributed Ceph clusters across physical data-centre infrastructure.
  • Manage MON, MGR, OSD, MDS and RGW services and client-to-cluster data-access patterns.
  • Add and remove nodes, migrate data and respond to disk and node failures.
  • Engineer CRUSH maps, data placement, replication and failure-domain strategies.
  • Diagnose degraded clusters, quorum issues, failed OSDs, uneven distribution and performance bottlenecks.
  • Control recovery, backfilling and rebalancing while protecting production workloads.
  • Investigate latency and throughput across the complete I/O path.
  • Operate SAN, NAS, DAS and object storage across Ceph and enterprise appliances.
  • Manage masking, zoning, RAID selection, snapshots and hardware maintenance.
  • Build capacity models and design backup, replication and disaster-recovery strategies.
  • Automate provisioning, configuration and upgrades using Bash, Python, Ansible, Puppet and Infrastructure as Code.
  • Monitor cluster health, capacity, latency, saturation and failure conditions using Prometheus, Grafana or comparable tooling.
  • Execute storage migrations and hardware refreshes without compromising availability.
  • Participate in a structured day and night on-call rotation.

What you bring

  • Deep expertise in either production Ceph or enterprise storage, with meaningful exposure to the other.
  • If Ceph is your primary domain, you have operated and troubleshot active production clusters, not only deployed them.
  • If you come from enterprise storage, you bring strong Linux and distributed-storage fundamentals and the aptitude to deepen your Ceph expertise.
  • Experience operating storage at meaningful scale, where performance and failure recovery require deliberate engineering.
  • Hands-on knowledge of Pure Storage, NetApp, Dell, Hitachi, HPE, IBM or similar platforms.
  • Experience with physical data centres and bare-metal infrastructure.
  • In-depth knowledge of the Linux storage stack, including the kernel block layer, LVM, multipathing, filesystems and performance diagnostics.
  • Strong understanding of FC, NVMe, iSCSI and NFS, plus block, file and object-storage workloads.
  • The ability to isolate problems across hardware, Linux, networking and distributed storage.
  • Automation experience with Bash and/or Python and configuration management such as Ansible or Puppet.
  • Experience with storage observability using Prometheus, Grafana or similar tooling.
  • Familiarity with AWS S3/EBS, KVM, VMware, GitOps or CI/CD is beneficial.
  • Willingness to take operational ownership and participate in on-call.

Why this is different

You will work where distributed-systems theory meets physical hardware and real production constraints. The infrastructure is not hidden behind a managed cloud service: the team owns its performance, resilience and failure modes.

Small engineering decisions can materially affect latency, recovery time, capacity and platform-wide availability. For storage engineers who enjoy working deep in the stack, this is an environment where technical judgement has visible impact.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Storage/System Engineers (Mid, Sr, Lead) – Ceph, Linux, Cloud & Enterprise Storage | 150+ PB scale | Mission-Critical Infrastructure at Global Scale FinTech
Storage/System Engineers (Mid, Sr, Lead) – Ceph, Linux, Cloud & Enterprise Storage | 150+ PB scale | Mission-Critical Infrastructure at Global Scale FinTech

Sprint and Partners • Illinois

On-site
USD 140,000 - 170,000
Senior Storage Engineer
Senior Storage Engineer

REALM • United States

Remote
USD 180,000 - 240,000
Storage Engineer
Storage Engineer

Net2Source (N2S) • Chicago (IL)

On-site
USD 82,656 - 89,544
Engineer, Storage and Data Protection
Engineer, Storage and Data Protection

AHEAD • New York (NY)

On-site
USD 120,000 - 180,000
Senior Staff Storage Engineer
Senior Staff Storage Engineer

DDN • Sacramento (CA)

On-site
USD 180,000 - 240,000
Early Career: IT Infrastructure - Storage Engineer
Early Career: IT Infrastructure - Storage Engineer

118-WW TMG MFG OPS • Dallas (TX)

On-site
USD 140,000 - 190,000
Lead Storage Engineer — Ceph, Linux & On-Prem
Lead Storage Engineer — Ceph, Linux & On-Prem

Sprint and Partners • San Francisco (CA)

On-site
USD 140,000 - 190,000
Senior Staff Engineer
Senior Staff Engineer

DDN • Santa Clara (CA)

Hybrid
USD 180,000 - 240,000
Storage Operations Engineer
Storage Operations Engineer

Webhosting • Northern (KY)

Hybrid
USD 80,000 - 100,000
Health insurance
401(k) plan with matching
Professional Development Reimbursement
+6
Senior System Administrator II (Ceph Engineer)
Senior System Administrator II (Ceph Engineer)

Adyen • Chicago (IL)

On-site
USD 140,000 - 210,000