Systems Resilience Engineer - Hybrid Cloud Storage

Qumulo

Seattle (WA)

Sur place

USD 140 000 - 210 000

Plein temps

14 jours+
Générateur de candidature

N’envoyez pas de CV générique — générez un CV et une lettre de motivation adaptés à ce poste précis.

Passez les filtres ATS

Avantages offerts par ce poste

Pre-IPO stock options
Flexible time-off policy
HSA and PPO health insurance
Dental and Vision insurance
401(k) plan
ORCA card or parking subsidy

Résumé du poste

Qumulo’s cloud data platform powers exabytes of data for thousands of customers across edge, core, and cloud. This Systems Resilience Engineer role focuses on breaking features, designing tests, and automating validation to ensure reliability across hardware and cloud environments.

You will design, implement, and maintain data-driven test plans, troubleshoot issues, and help define release quality bar. Open metrics and monitoring culture are emphasized.

Qualifications

  • 3+ years building and operating automated testing, validation, and/or certification for complex software systems.
  • Strong programming ability in C; experience with distributed file systems is a plus.
  • Breaker mindset with edge-case discovery and proactive testing.
  • Proven track record of building tests themselves, not just following plans.
  • Hands-on across on-prem and cloud (AWS, GCP, or Azure).
  • Fluency in Linux and Python.
  • Experience with orchestration tools (Ansible, Terraform), containers, and Kubernetes.
  • Solid understanding of networks; storage or protocol experience a plus.

Responsabilités

  • Design and operationalize testing for new features, including scale testing and breakage analysis.
  • Automate manual, repetitive testing using Python and in-house frameworks on Jenkins and Argo.
  • Build a data-driven plan for tests, including scheduling and reruns.
  • Troubleshoot build and test failures across VM instances and hardware.
  • Read logs to distinguish test vs infrastructure vs real bug issues.
  • Set up monitoring and alerting with OpenMetrics, Grafana, InfluxDB, and Prometheus.
  • Help set the quality bar for releases and participate in release decisions.
  • Take part in an on-call rotation for the team's systems.

Connaissances

Automated testing
C programming
Linux
Python
Cloud platforms
Jenkins
Argo
Kubernetes
Ansible
Terraform
Networking fundamentals
Data-driven testing
Distributed file systems

Outils

Jenkins
Argo
Prometheus
Grafana

Description du poste

Qumulo's cloud data platform manages exabytes of the world's most demanding data, unifying files, objects, and every workload across edge, core, and cloud.

About The Role

You’ll be one of the first hires on a team with a single mandate: find out how Qumulo breaks before our customers do. The platform manages exabytes of data for more than 1,100 customers across on-prem and every major cloud, and these are mission-critical workloads where a missed edge case becomes a customer's bad day.

This is a Systems Resilience Engineer role for the engineer who thinks like a breaker. You'll put on the customer's hat, work out how a feature will really get used, and design the tests that push it past its limits across hardware and cloud. You'll automate the testing our principal engineers run by hand today, and decide what gets tested, how often, and why. You'll help build this function from the ground up, including where we set the quality bar and which builds are good enough to ship.

As a Systems Resilience Engineer At Qumulo You Will
  • Design and operationalize testing for new features: work out how customers will actually use them, how to scale-test them, and how to break them
  • Automate the manual, repetitive testing our principal engineers run by hand today, using Python and our in-house frameworks on Jenkins and Argo
  • Build a data-driven plan for which tests run, how often, and why, plus the framework to schedule and rerun them
  • Troubleshoot build and test failures across VM instances and Qumulo-qualified hardware, from compile-time errors to integration failures
  • Read cluster output and C error logs to tell a test problem from an infrastructure problem from a real bug
  • Set up monitoring and alerting so problems surface early (we use OpenMetrics, Grafana, InfluxDB, and Prometheus alongside home-grown tooling)
  • Help set the quality bar for releases, including a real say in what ships
  • Take part in an on-call rotation for the systems your team owns
Our Ideal Candidate Will Have
  • 3+ years building and operating automated testing, validation, and/or certification for complex software systems
  • Strong programming ability in C. Experience with Qumulo’s distributed file system, or parallel filesystems, would be a major plus
  • A real breaker's instinct. You go looking for edge cases and ask "what happens if I do this?" before anyone asks you to
  • A track record of building tests yourself, not just running test plans handed to you
  • Hands-on experience across both on-premises infrastructure and cloud (AWS, GCP, or Azure), with a real grasp of where each one's limits are
  • Working fluency in Linux (we run Ubuntu) and Python
  • A data-driven approach to deciding what to test and how often
  • Experience with orchestration tools (Ansible, Terraform), containers, and Kubernetes
  • Solid understanding of networks (routing, firewalls, security inspection devices, switch configuration) a plus
  • Storage (IOPS, Latency, read/write patterns) or protocol experience (NFS, SMB, S3, ) a strong plus

The annual pay range for the role is USD $140,000 - $210,000. Individual pay depends on various factors, such as role level, relevant experience, and skills.

Benefits & Perks
  • Pre-IPO stock options
  • Flexible time-off policy
  • HSA and PPO health insurance options
  • Dental and Vision insurance
  • 401(k) plan
  • Choice of an ORCA card or parking subsidy
About Qumulo

Built for the most demanding enterprise workloads, from AI and HPC simulations to Splunk Observability, genomics and PACS medical imaging, geospatial datasets, media editing and rendering, and video surveillance, Qumulo unlocks the full power of an organization’s information. Our platform unifies file and object storage across data centers, edge, and public clouds, enabling efficient, accelerated computing and extending the reach of data unbound by protocol or transport limitations.

With more than 1,100 customers and exabytes of data under management, Qumulo powers mission-critical workloads anywhere real-time access to massive file datasets is non-negotiable. Qumulo delivers radical simplicity, hardware freedom, exceptional customer support, and a true hybrid-cloud architecture.

Our Values

At Qumulo, we are building an open and collaborative culture where people can do their best work with customers as our magnetic field. We act as owners, we share by default, we are data driven and experimental and as an inclusive workplace, we encourage and celebrate multiple points of view. As part of our culture we believe diversity drives innovation.

Equal Opportunity Employer

Qumulo is an Equal Opportunity Employer. Qualified applicants will receive consideration for employment without regard to race, color, religion, sex, age, disability, military status, national origin, or any other characteristic protected under federal, state, or applicable local law. For more information on Qumulo's Applicant Privacy Policy, please visit: https://qumulo.com/applicant-employee-privacy-notice

Obtenez votre examen gratuit et confidentiel de votre CV.

ou faites glisser et déposez votre fichier ici.

Similar jobs

Postes similaires à comparer

Systems Resilience Engineer - Hybrid Cloud Storage
Systems Resilience Engineer - Hybrid Cloud Storage

Madrona Venture Labs • Seattle (WA)

Sur place
USD 140 000 - 210 000
Pre-IPO stock options
Flexible time-off policy
HSA and PPO health insurance options
+3
Senior Systems Engineer (Enterprise Storage) - PA or NJ
Senior Systems Engineer (Enterprise Storage) - PA or NJ

Madrona Venture Labs • Pennsylvania

Sur place
USD 250 000 - 275 000
Pre-IPO stock options
Flexible time-off policy
Health insurance options
+3
Senior Member of Technical Staff
Senior Member of Technical Staff

Qumulo • Seattle (WA)

Hybride
USD 200 000 - 300 000
Software Development Engineer (New Grad / Entry Level)
Software Development Engineer (New Grad / Entry Level)

Qumulo • Seattle (WA)

Hybride
USD 110 000 - 140 000
Pre-IPO stock options
Flexible time-off policy
HSA and PPO health insurance options
+3
Senior Member of Technical Staff
Senior Member of Technical Staff

Madrona Venture Labs • Seattle (WA)

Sur place
USD 200 000 - 300 000
Stock options
Health insurance
Dental & Vision insurance
+4
Customer Success Engineer II (Cloud / Storage / Compute / Identity) - UK Remote
Customer Success Engineer II (Cloud / Storage / Compute / Identity) - UK Remote

Qumulo • États-Unis

À distance
USD 90 000 - 120 000
Remote Customer Success Engineer II (Cloud / Storage / Compute / Identity) - UK Remote
Remote Customer Success Engineer II (Cloud / Storage / Compute / Identity) - UK Remote

Qumulo, Inc. • États-Unis

À distance
USD 90 000 - 130 000
Software Development Engineer (Seattle) - Internship 2027
Software Development Engineer (Seattle) - Internship 2027

Qumulo • Seattle (WA), Northern (KY)

Hybride
USD 1 037 000 - 1 128 000
Relocation support
Travel support
Transit support
Senior Systems Engineer (Federal) - DC or Virginia
Senior Systems Engineer (Federal) - DC or Virginia

Madrona Venture Labs • Virginia (MN)

Sur place
USD 260 000 - 270 000
Pre-IPO stock options
Flexible time-off policy
HSA and PPO health insurance options
+3
Senior Customer Success Manager - Data / Storage / Cloud (Central or East)
Senior Customer Success Manager - Data / Storage / Cloud (Central or East)

Qumulo • Mississippi

Sur place
USD 125 000 - 180 000
Pre-IPO stock options
Flexible time-off policy
HSA and PPO health insurance options
+3