Staff Engineer, Lustre

Ddn

Santa Clara (CA)

On-site

USD 190,000 - 260,000

Full time

5 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Ddn is seeking a Staff Engineer with 10+ years of experience in distributed storage and Linux-based systems engineering. The role is hands-on and focuses on design, debugging, performance, and operational excellence across LustreFS and related components.

The ideal candidate will drive complex investigations, collaborate across engineering, QE, support and release teams, and use AI-assisted tools to speed triage, logging, and feature design. Strong subsystem depth is required.

Qualifications

  • 10+ years in systems software or distributed storage.
  • Strong LustreFS experience and subsystem depth.
  • Proficient in C programming and Linux debugging.
  • Experience across client, server, network and storage layers.

Responsibilities

  • Hands-on development and debugging of LustreFS features and fixes.
  • Investigate scale-related defects and perform root-cause analysis.
  • Contribute to performance tuning and reliability improvements.
  • Participate in design reviews and subsystem discussions.
  • Collaborate with QE, support and release teams for diagnostics.

Skills

Distributed systems
Linux kernel
C programming
Performance analysis
AI-assisted tooling
Code reviews

Education

BS in Computer Science

Tools

LustreFS
LNet
RDMA
NVMe tooling

Job description

We are seeking a Staff Engineer with 10+ years of experience in distributed storage and Linux-based systems engineering. This is a hands‑on senior technical role focused on design, debugging, performance, and operational excellence across LustreFS and adjacent stack components. The ideal candidate brings strong expertise in one or more Lustre subsystems, can independently drive complex investigations, and collaborates effectively across engineering, QE, support and release teams. Engineers who are comfortable using AI to accelerate triage, debugging, code comprehension and new feature design will be especially valuable.

Key Responsibilities
  • Design, develop and debug LustreFS features, fixes and enhancements across relevant subsystems such as llite, MDS/MDT, OSS/OST, LDLM and LNet.

  • Investigate customer and scale-related defects, drive root‑cause analysis and implement high‑quality fixes with strong attention to correctness and maintainability.

  • Contribute to performance tuning, failure analysis and reliability improvements for large-scale Lustre deployments.

  • Participate actively in code reviews, design reviews and subsystem discussions, bringing rigor to testing and operational readiness.

  • Work closely with QE and support to reproduce issues, improve diagnostic data quality and increase coverage for high‑risk failure scenarios.

  • Help document subsystem behavior, debugging approaches, known failure patterns and operational best practices.

  • Use AI‑assisted tools where appropriate to speed up issue triage, summarize logs, improve code understanding and capture reusable lessons learned.

Required Qualifications
  • 10+ years of experience in systems software, distributed systems, storage, Linux kernel or filesystem engineering.

  • Strong experience in LustreFS development, support or performance engineering with depth in at least one major subsystem.

  • Strong C programming and Linux systems debugging skills.

  • Working knowledge of Linux kernel internals, filesystem semantics, networking and performance analysis.

  • Experience with LNet and/or high-performance transports such as RDMA, InfiniBand, RoCE or TCP-based storage networking.

  • Ability to debug and resolve issues spanning multiple layers including client, server, network and backend storage.

  • Strong collaboration skills and the ability to work across functions in a fast‑moving engineering environment.

Preferred Skills
  • Experience in HPC, AI infrastructure or large‑scale parallel storage environments.

  • Exposure to metadata-heavy and throughput-heavy workload characterization and tuning.

  • Familiarity with ZFS, ldiskfs, NVMe-backed storage and related observability / performance tooling.

  • Experience creating test plans, reproducer frameworks, runbooks or diagnostic automation.

  • Comfort using AI tools to accelerate debugging, code reviews, triage, documentation and early‑stage design ideation.

  • Experience mentoring junior engineers or leading focused technical efforts within a subsystem.

What You Will Work On
  • Hands‑on development and debugging of LustreFS defects, performance issues and subsystem enhancements.

  • Customer‑facing and scale‑related issue investigation across llite, metadata, object storage, LNet and transport layers.

  • Collaborative design and implementation of reliability, observability and serviceability improvements.

  • Reviewing and validating fixes through targeted tests, failure injection, log analysis and performance characterization.

  • Using AI-assisted workflows to accelerate triage, debug loops, code understanding and documentation quality.

  • Contributing to team redundancy by strengthening documentation, code review quality and subsystem knowledge sharing.

Why This Role Matters

This role is central to building durable engineering redundancy in LustreFS: expanding deep subsystem ownership, reducing concentration risk, and accelerating next‑generation delivery through strong engineering fundamentals and AI-enabled execution.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Sr/Staff Lustre Engineer
Sr/Staff Lustre Engineer

DDN • Santa Clara (CA)

On-site
USD 150,000 - 230,000
Senior LustreFS Engineer: Performance & AI Debugging
Senior LustreFS Engineer: Performance & AI Debugging

Ddn • Santa Clara (CA)

On-site
USD 190,000 - 260,000
Senior Lustre Architect & Performance Lead
Senior Lustre Architect & Performance Lead

Data Direct Networks • California (MO)

On-site
USD 140,000 - 210,000
Highly Competitive Vacation Plans
Paid Holidays
Bonus Programs
+3
Senior Lustre Engineer — Kernel & Performance Architect
Senior Lustre Engineer — Kernel & Performance Architect

DDN • Santa Clara (CA)

On-site
USD 150,000 - 230,000
Senior Lustre Engineer & Tech Lead
Senior Lustre Engineer & Tech Lead

NetApp • California (MO)

Hybrid
USD 170,000 - 253,000
Lustre Software Engineer
Lustre Software Engineer

NetApp • California (MO)

Hybrid
USD 170,000 - 253,000
File Systems Engineer - HPC Storage & Lustre Expert
File Systems Engineer - HPC Storage & Lustre Expert

Hewlett Packard Enterprise Development LP • BLOOMINGTON (MN)

On-site
USD 106,000 - 243,000
File Systems Engineer
File Systems Engineer

Hewlett Packard Enterprise Company • Bloomington (IN)

Hybrid
USD 110,000 - 160,000
Senior/Staff Fuse Developer
Senior/Staff Fuse Developer

DDN • United States

Remote
USD 120,000 - 180,000
HPC Lustre Storage Engineer
HPC Lustre Storage Engineer

Hewlett Packard Enterprise • BLOOMINGTON (MN)

Hybrid
USD 106,000 - 243,000
Health & Wellbeing program
Professional development
Inclusive culture