Storage/System Engineers (Mid, Sr, Lead) – Ceph, Linux, Cloud & Enterprise Storage | 150+ PB scale | Mission-Critical Infrastructure at Global Scale FinTech

Sprint and Partners

Illinois

On-site

USD 140,000 - 170,000

Full time

6 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Sprint and Partners is seeking a senior storage engineer to design, operate and evolve Ceph-based infrastructure across on-premises data centers. You will own the storage stack from hardware to software layers and help ensure data integrity, performance and availability for a global fintech platform.

As part of a small, globally-distributed team, you will troubleshoot latency, plan capacity, and implement automation in Bash and Python, with on-call responsibilities and collaboration with

Qualifications

  • Deep expertise in production Ceph or enterprise storage.
  • Experience operating storage at meaningful scale with failure recovery.
  • Hands-on knowledge of physical data centres and bare-metal infra.
  • Familiarity with FC, NVMe, iSCSI and NFS workloads; block, file and object storage.

Responsibilities

  • Design, operate and evolve storage infrastructure supporting a global fintech platform.
  • Manage MON, MGR, OSD, MDS and RGW services and data-access patterns.
  • Add/remove nodes, migrate data and respond to disk and node failures.
  • Engineer CRUSH maps, data placement and failure-domain strategies.
  • Diagnose degraded clusters, quorum issues and performance bottlenecks.
  • Automate provisioning, configuration and upgrades with Bash, Python, Ansible or Puppet.

Skills

Ceph expertise
Linux fundamentals
Distributed storage
Automation Bash/Python
Prometheus Grafana
On-call

Tools

Pure Storage
NetApp
Dell
Hitachi
HPE
IBM

Job description

This global fintech operates business-critical infrastructure at significant scale (150 PB+). Its storage platform supports workloads ranging from high-throughput, low-latency transactional databases to petabyte-scale object stores, making data integrity and availability fundamental to the product.

This is not a conventional public-cloud role. The core platform runs on on-premises, bare-metal infrastructure, and engineers own the full storage stack: hardware, Linux, distributed storage, automation, observability, capacity and failure recovery. The company is moving to Cloud which will be an exciting process as well.

The environment combines Ceph with enterprise platforms such as Pure Storage, NetApp, Dell, Hitachi and HPE. You may be strongest in either software-defined or enterprise storage, provided you have strong fundamentals and can develop further in the other domain.

Engineers currently focused on cloud infrastructure are also relevant if they have previous hands-on bare-metal or on-premises experience.

The role

  • Design, operate and evolve storage infrastructure supporting a global financial technology platform.
  • Reason about data placement, replication, quorum, failure domains, recovery behaviour, latency, throughput and capacity.
  • Trace complex problems across disks, controllers, networks, the Linux storage stack and Ceph.
  • Become the first Chicago-based engineer in an established Storage team of approximately six engineers, currently centred in Amsterdam.
  • Collaborate with database, networking, application and Platform Engineering teams.

What you'll work on

  • Design and operate distributed Ceph clusters across physical data-centre infrastructure.
  • Manage MON, MGR, OSD, MDS and RGW services and client-to-cluster data-access patterns.
  • Add and remove nodes, migrate data and respond to disk and node failures.
  • Engineer CRUSH maps, data placement, replication and failure-domain strategies.
  • Diagnose degraded clusters, quorum issues, failed OSDs, uneven distribution and performance bottlenecks.
  • Control recovery, backfilling and rebalancing while protecting production workloads.
  • Investigate latency and throughput across the complete I/O path.
  • Operate SAN, NAS, DAS and object storage across Ceph and enterprise appliances.
  • Manage masking, zoning, RAID selection, snapshots and hardware maintenance.
  • Build capacity models and design backup, replication and disaster-recovery strategies.
  • Automate provisioning, configuration and upgrades using Bash, Python, Ansible, Puppet and Infrastructure as Code.
  • Monitor cluster health, capacity, latency, saturation and failure conditions using Prometheus, Grafana or comparable tooling.
  • Execute storage migrations and hardware refreshes without compromising availability.
  • Participate in a structured day and night on-call rotation.

What you bring

  • Deep expertise in either production Ceph or enterprise storage, with meaningful exposure to the other.
  • If Ceph is your primary domain, you have operated and troubleshot active production clusters, not only deployed them.
  • If you come from enterprise storage, you bring strong Linux and distributed-storage fundamentals and the aptitude to deepen your Ceph expertise.
  • Experience operating storage at meaningful scale, where performance and failure recovery require deliberate engineering.
  • Hands-on knowledge of Pure Storage, NetApp, Dell, Hitachi, HPE, IBM or similar platforms.
  • Experience with physical data centres and bare-metal infrastructure.
  • In-depth knowledge of the Linux storage stack
  • Strong understanding of FC, NVMe, iSCSI and NFS, plus block, file and object-storage workloads.
  • The ability to isolate problems across hardware, Linux, networking and distributed storage.
  • Automation experience with Bash and/or Python and configuration management such as Ansible or Puppet.
  • Experience with storage observability using Prometheus, Grafana or similar tooling.
  • Familiarity with AWS S3/EBS, KVM, VMware, GitOps or CI/CD is beneficial.
  • Willingness to take operational ownership and participate in on-call.

Why this is different

You will work where distributed-systems theory meets physical hardware and real production constraints. The infrastructure is not hidden behind a managed cloud service: the team owns its performance, resilience and failure modes.

Small engineering decisions can materially affect latency, recovery time, capacity and platform-wide availability. For storage engineers who enjoy working deep in the stack, this is an environment where technical judgement has visible impact.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Storage/System Engineers (Mid, Sr, Lead) – Ceph, Linux, Cloud & Enterprise Storage | 150+ PB scale | Mission-Critical Infrastructure at Global Scale FinTech
Storage/System Engineers (Mid, Sr, Lead) – Ceph, Linux, Cloud & Enterprise Storage | 150+ PB scale | Mission-Critical Infrastructure at Global Scale FinTech

Sprint and Partners • San Francisco (CA)

On-site
USD 140,000 - 190,000
Senior Storage Engineer
Senior Storage Engineer

REALM • United States

Remote
USD 180,000 - 240,000
Engineer, Storage and Data Protection
Engineer, Storage and Data Protection

AHEAD • New York (NY)

On-site
USD 120,000 - 180,000
Storage Engineer
Storage Engineer

Net2Source (N2S) • Chicago (IL)

On-site
Senior Staff Storage Engineer
Senior Staff Storage Engineer

DDN • Sacramento (CA)

On-site
USD 180,000 - 240,000
Early Career: IT Infrastructure - Storage Engineer
Early Career: IT Infrastructure - Storage Engineer

118-WW TMG MFG OPS • Dallas (TX)

On-site
USD 140,000 - 190,000
Lead Storage Engineer — Ceph, Linux & On-Prem
Lead Storage Engineer — Ceph, Linux & On-Prem

Sprint and Partners • San Francisco (CA)

On-site
USD 140,000 - 190,000
Senior Staff Engineer
Senior Staff Engineer

DDN • Santa Clara (CA)

Hybrid
USD 180,000 - 240,000
Storage Operations Engineer
Storage Operations Engineer

Webhosting • Northern (KY)

Hybrid
USD 80,000 - 100,000
Health insurance
401(k) plan with matching
Professional Development Reimbursement
+6
Senior System Administrator II (Ceph Engineer)
Senior System Administrator II (Ceph Engineer)

Adyen • Chicago (IL)

On-site
USD 140,000 - 210,000