Systems Resilience Engineer - Hybrid Cloud Storage

Qumulo

Seattle (WA)

On-site

USD 140,000 - 210,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Pre-IPO stock options
Flexible time-off policy
HSA and PPO health insurance
Dental and Vision insurance
401(k) plan
ORCA card or parking subsidy

Job summary

Qumulo’s cloud data platform powers exabytes of data for thousands of customers across edge, core, and cloud. This Systems Resilience Engineer role focuses on breaking features, designing tests, and automating validation to ensure reliability across hardware and cloud environments.

You will design, implement, and maintain data-driven test plans, troubleshoot issues, and help define release quality bar. Open metrics and monitoring culture are emphasized.

Qualifications

  • 3+ years building and operating automated testing, validation, and/or certification for complex software systems.
  • Strong programming ability in C; experience with distributed file systems is a plus.
  • Breaker mindset with edge-case discovery and proactive testing.
  • Proven track record of building tests themselves, not just following plans.
  • Hands-on across on-prem and cloud (AWS, GCP, or Azure).
  • Fluency in Linux and Python.
  • Experience with orchestration tools (Ansible, Terraform), containers, and Kubernetes.
  • Solid understanding of networks; storage or protocol experience a plus.

Responsibilities

  • Design and operationalize testing for new features, including scale testing and breakage analysis.
  • Automate manual, repetitive testing using Python and in-house frameworks on Jenkins and Argo.
  • Build a data-driven plan for tests, including scheduling and reruns.
  • Troubleshoot build and test failures across VM instances and hardware.
  • Read logs to distinguish test vs infrastructure vs real bug issues.
  • Set up monitoring and alerting with OpenMetrics, Grafana, InfluxDB, and Prometheus.
  • Help set the quality bar for releases and participate in release decisions.
  • Take part in an on-call rotation for the team's systems.

Skills

Automated testing
C programming
Linux
Python
Cloud platforms
Jenkins
Argo
Kubernetes
Ansible
Terraform
Networking fundamentals
Data-driven testing
Distributed file systems

Tools

Jenkins
Argo
Prometheus
Grafana

Job description

Qumulo's cloud data platform manages exabytes of the world's most demanding data, unifying files, objects, and every workload across edge, core, and cloud.

About The Role

You’ll be one of the first hires on a team with a single mandate: find out how Qumulo breaks before our customers do. The platform manages exabytes of data for more than 1,100 customers across on-prem and every major cloud, and these are mission-critical workloads where a missed edge case becomes a customer's bad day.

This is a Systems Resilience Engineer role for the engineer who thinks like a breaker. You'll put on the customer's hat, work out how a feature will really get used, and design the tests that push it past its limits across hardware and cloud. You'll automate the testing our principal engineers run by hand today, and decide what gets tested, how often, and why. You'll help build this function from the ground up, including where we set the quality bar and which builds are good enough to ship.

As a Systems Resilience Engineer At Qumulo You Will
  • Design and operationalize testing for new features: work out how customers will actually use them, how to scale-test them, and how to break them
  • Automate the manual, repetitive testing our principal engineers run by hand today, using Python and our in-house frameworks on Jenkins and Argo
  • Build a data-driven plan for which tests run, how often, and why, plus the framework to schedule and rerun them
  • Troubleshoot build and test failures across VM instances and Qumulo-qualified hardware, from compile-time errors to integration failures
  • Read cluster output and C error logs to tell a test problem from an infrastructure problem from a real bug
  • Set up monitoring and alerting so problems surface early (we use OpenMetrics, Grafana, InfluxDB, and Prometheus alongside home-grown tooling)
  • Help set the quality bar for releases, including a real say in what ships
  • Take part in an on-call rotation for the systems your team owns
Our Ideal Candidate Will Have
  • 3+ years building and operating automated testing, validation, and/or certification for complex software systems
  • Strong programming ability in C. Experience with Qumulo’s distributed file system, or parallel filesystems, would be a major plus
  • A real breaker's instinct. You go looking for edge cases and ask "what happens if I do this?" before anyone asks you to
  • A track record of building tests yourself, not just running test plans handed to you
  • Hands-on experience across both on-premises infrastructure and cloud (AWS, GCP, or Azure), with a real grasp of where each one's limits are
  • Working fluency in Linux (we run Ubuntu) and Python
  • A data-driven approach to deciding what to test and how often
  • Experience with orchestration tools (Ansible, Terraform), containers, and Kubernetes
  • Solid understanding of networks (routing, firewalls, security inspection devices, switch configuration) a plus
  • Storage (IOPS, Latency, read/write patterns) or protocol experience (NFS, SMB, S3, ) a strong plus

The annual pay range for the role is USD $140,000 - $210,000. Individual pay depends on various factors, such as role level, relevant experience, and skills.

Benefits & Perks
  • Pre-IPO stock options
  • Flexible time-off policy
  • HSA and PPO health insurance options
  • Dental and Vision insurance
  • 401(k) plan
  • Choice of an ORCA card or parking subsidy
About Qumulo

Built for the most demanding enterprise workloads, from AI and HPC simulations to Splunk Observability, genomics and PACS medical imaging, geospatial datasets, media editing and rendering, and video surveillance, Qumulo unlocks the full power of an organization’s information. Our platform unifies file and object storage across data centers, edge, and public clouds, enabling efficient, accelerated computing and extending the reach of data unbound by protocol or transport limitations.

With more than 1,100 customers and exabytes of data under management, Qumulo powers mission-critical workloads anywhere real-time access to massive file datasets is non-negotiable. Qumulo delivers radical simplicity, hardware freedom, exceptional customer support, and a true hybrid-cloud architecture.

Our Values

At Qumulo, we are building an open and collaborative culture where people can do their best work with customers as our magnetic field. We act as owners, we share by default, we are data driven and experimental and as an inclusive workplace, we encourage and celebrate multiple points of view. As part of our culture we believe diversity drives innovation.

Equal Opportunity Employer

Qumulo is an Equal Opportunity Employer. Qualified applicants will receive consideration for employment without regard to race, color, religion, sex, age, disability, military status, national origin, or any other characteristic protected under federal, state, or applicable local law. For more information on Qumulo's Applicant Privacy Policy, please visit: https://qumulo.com/applicant-employee-privacy-notice

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Systems Resilience Engineer - Hybrid Cloud Storage
Systems Resilience Engineer - Hybrid Cloud Storage

Madrona Venture Labs • Seattle (WA)

On-site
USD 140,000 - 210,000
Pre-IPO stock options
Flexible time-off policy
HSA and PPO health insurance options
+3
Site Reliability Engineer (SRE) - Hybrid Cloud Storage
Site Reliability Engineer (SRE) - Hybrid Cloud Storage

Qumulo • Seattle (WA)

On-site
USD 140,000 - 210,000
Pre-IPO stock options
Flexible time-off policy
HSA and PPO health insurance options
+3
Senior Systems Engineer (Enterprise Storage) - PA or NJ
Senior Systems Engineer (Enterprise Storage) - PA or NJ

Madrona Venture Labs • Pennsylvania

On-site
USD 250,000 - 275,000
Pre-IPO stock options
Flexible time-off policy
Health insurance options
+3
Software Engineer, Core Data Services
Software Engineer, Core Data Services

Qumulo • Seattle (WA)

On-site
USD 140,000 - 190,000
Pre-IPO stock options
Flexible time-off policy
HSA and PPO health insurance options
+3
Senior Systems Administrator
Senior Systems Administrator

Qumulo • Seattle (WA)

On-site
USD 120,000 - 150,000
Pre-IPO stock options
Flexible time-off policy
HSA and PPO health insurance options
+3
Senior Systems Administrator
Senior Systems Administrator

Madrona Venture Labs • Seattle (WA)

On-site
USD 120,000 - 150,000
Pre-IPO stock options
Flexible time-off policy
HSA and PPO health insurance options
+3
Pre-Sales Solutions Architect - PNW
Pre-Sales Solutions Architect - PNW

Socket.dev • Seattle (WA)

On-site
USD 175,000 - 265,000
Pre-IPO stock options
Flexible time-off policy
HSA and PPO health insurance options
+3
Pre-Sales Solutions Architect (LA) - Media & Entertainment - Enterprise Data Storage / Hybrid Cloud
Pre-Sales Solutions Architect (LA) - Media & Entertainment - Enterprise Data Storage / Hybrid Cloud

Qumulo • Los Angeles (CA)

Hybrid
USD 175,000 - 265,000
Pre-IPO stock options
Flexible time-off policy
Health insurance options
+2
Technical Program Manager (TPM) - Hybrid Cloud Storage
Technical Program Manager (TPM) - Hybrid Cloud Storage

Qumulo • Seattle (WA)

Hybrid
USD 125,000 - 200,000
Pre-IPO stock options
Flexible time-off policy
HSA and PPO health insurance options
+3
Senior Systems Engineer (Federal) - DC or Virginia
Senior Systems Engineer (Federal) - DC or Virginia

Madrona Venture Labs • Virginia (MN)

On-site
USD 260,000 - 270,000
Pre-IPO stock options
Flexible time-off policy
HSA and PPO health insurance options
+3