Senior Storage Engineer for AI GPU Clusters - DGX Cloud

2100 NVIDIA USA

Santa Clara (CA)

On-site

USD 224,000 - 431,250

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity
Benefits

Job summary

NVIDIA DGXC Storage is seeking a hands‑on Storage Software Engineer to lead development on open‑source parallel and distributed file systems, ensuring the reliability, durability, and performance of our largest GPU clusters.

You will contribute to production code, diagnose complex storage issues, and set standards for configuration and tuning across the fleet. A deep Linux storage background and strong C/C++/Rust/Go skills are essential.

Qualifications

  • BS, MS, or PhD in Computer Science, Electrical Engineering, or a related field — or equivalent experience.
  • Over 12 years of direct experience in storage software engineering, including extensive involvement with a high‑performance parallel or distributed file system handling multi‑petabyte scale.
  • Contributions to open‑source projects involving a distributed or parallel file system. You write and review production code, examine file system, kernel, NVMe‑oF, or SPDK source to identify bugs, and personally conduct scale tests or recovery drills instead of assigning them to others.
  • Experience diagnosing and resolving storage problems in extensive GPU or HPC clusters, including analysis of I/O and metadata performance.
  • Strong proficiency in at least one systems language (C, C++, Rust, or Go) and proficiency in Python; comfortable in the Linux kernel storage and networking stacks (block layer, RDMA / RoCE / InfiniBand, NVMe, page cache, VFS, multipath).
  • Working knowledge of object storage (S3 / Swift‑class) and block storage (NVMe‑oF, iSCSI).
  • Strong written and verbal communication; capable of clarifying complex technical trade‑offs to engineers, SREs, vendors, and internal customers.
  • Comfort operating in a 24/7 production environment where storage incidents directly impact GPU availability, with a security‑first approach baked into every build.
  • 100% hands‑on engineering. You write and review production code, read file system, kernel, NVMe‑oF, or SPDK source to chase bugs, and run scale tests or recovery drills yourself rather than delegating.

Responsibilities

  • Contribute to open‑source file systems. Contribute code to open‑source parallel and distributed file systems, and distributed object storage. Upstream fixes and features, and engage directly with the upstream communities and maintainers.
  • Serve as a hands‑on storage software lead. Write and review production code yourself, and read kernel, NFS, NVMe‑oF, or SPDK source when a bug requires it. Make the final technical calls on storage deliveries against measurable targets.
  • Triage and troubleshoot at scale. Triage, troubleshoot, and root‑cause large, complex storage issues across very large GPU clusters (tens of thousands of GPUs) — I/O and metadata performance, data corruption, and recovery.
  • Validate architecture and capabilities. Validate storage architecture, capabilities, performance, and durability. Run scale tests, benchmarks, and recovery drills, and qualify new builds against measurable performance and durability targets.
  • Recommend configuration, tuning, and guidelines. Define and recommend configuration, tuning, and operational best practices for high‑performance file systems on GPU infrastructure, and help operators and internal customers apply them.
  • Partner broadly. Work with training, inference, and accelerated‑computing teams, site‑reliability and operations, networking, and security, and collaborate with cloud providers, neocloud operators, and storage vendors on a common architecture.
  • Work AI‑first. Use modern AI coding and agentic tools day‑to‑day to accelerate building, debugging, validation, and operations.

Skills

C/C++/Rust/Go
Python
Kubernetes
Linux kernel knowledge

Education

BS/MS/PhD in CS/EE or related

Tools

SPDK
NVMe-oF
Kubernetes CSI

Job description

NVIDIA DGXC Storage is seeking a hands‑on Storage Software Engineer to lead development on open‑source parallel and distributed file systems, ensuring the reliability, durability, and performance of our largest GPU clusters.

You will contribute to production code, diagnose complex storage issues, and set standards for configuration and tuning across the fleet. A deep Linux storage background and strong C/C++/Rust/Go skills are essential.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI Storage Software Engineer — Scale GPU Clusters
Senior AI Storage Software Engineer — Scale GPU Clusters

Nvidia Corporation • Santa Clara (CA)

On-site
USD 224,000 - 432,000
Equity
Benefits
Senior Storage Systems Engineer for AI Data Infrastructure
Senior Storage Systems Engineer for AI Data Infrastructure

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 152,000 - 288,000
Equity compensation
Benefits
Senior Storage Engineer — Cloud-Native AI Data Infrastructure
Senior Storage Engineer — Cloud-Native AI Data Infrastructure

NVIDIA • California (MO)

On-site
USD 184,000 - 288,000
Equity
Benefits
Senior Storage Systems Engineer for AI & Cloud Data
Senior Storage Systems Engineer for AI & Cloud Data

NVIDIA • Town of Texas (WI)

On-site
USD 152,000 - 288,000
Equity
Benefits
Senior AI Storage Systems Engineer
Senior AI Storage Systems Engineer

NVIDIA • United States

On-site
USD 152,000 - 288,000
Equity
Benefits
Senior Storage Systems Engineer: Kubernetes & Automation (Remote)
Senior Storage Systems Engineer: Kubernetes & Automation (Remote)

NVIDIA Corporation • United States

On-site
USD 208,000 - 414,000
Equity
Benefits package
Senior Storage Engineer for AI & Kubernetes
Senior Storage Engineer for AI & Kubernetes

United States Digital Space LLC • United States

Remote
USD 140,000 - 200,000
Senior Storage DevOps Engineer for GPU AI Platforms
Senior Storage DevOps Engineer for GPU AI Platforms

United States Digital Space LLC • United States

Remote
USD 120,000 - 190,000
Competitive compensation
Strong benefits plan
Conference attendance
+1
Senior Storage Production Engineer: Low-Latency AI Cloud
Senior Storage Production Engineer: Low-Latency AI Cloud

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 176,000 - 276,000
Equity opportunities
Comprehensive benefits package
Senior Storage Software Engineer - DGX Cloud
Senior Storage Software Engineer - DGX Cloud

Nvidia Corporation • Santa Clara (CA)

On-site
USD 224,000 - 432,000
Equity
Benefits