Remote: Head of Storage Production Engineering

NVIDIA Corporation

Santa Clara (CA)

Hybrid

USD 272,000 - 431,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

NVIDIA is seeking a leader for Storage Production Engineering to design, deploy, and operate large-scale storage systems that power AI and HPC workloads. You will build a collaborative, learning-driven team and guide data pipelines and storage technologies across the platform.

You will own capacity planning, high availability, and incident response while partnering with engineering, DevOps, and AI/ML teams to improve performance and reliability.

Qualifications

  • BS or MS in Computer Science, Storage Systems, or related field.
  • 12+ years in large-scale storage architecture, operations, production engineering, or infrastructure.
  • 6+ years of people management or technical leadership in storage, infrastructure, or site reliability teams.
  • Experience managing infrastructure operations including on-call rotations, incident response, maintenance, troubleshooting, and optimization of production systems with SLOs/KPIs.
  • Hands-on with parallel file systems (Lustre/GPFS), distributed storage (Ceph/MinIO), and S3-compatible/NAS platforms (NetApp, Pure Storage).
  • Strong knowledge of block, file, and object storage and HA design.
  • Experience with storage networking and protocols (NFS, SMB, iSCSI, Fibre Channel, RDMA, NVMe-oF).
  • IaC tooling: Terraform, Ansible, Puppet.
  • Knowledge of monitoring/observability tools (Prometheus, InfluxDB, Elastic).

Responsibilities

  • Lead and coach a team of Storage Production Engineers.
  • Design, deploy, and improve large-scale storage systems (distributed storage, parallel file systems, object storage).
  • Use automation, monitoring, and analytics to improve reliability and efficiency of storage services.
  • Own capacity planning, data lifecycle management, cost awareness, and high availability/DR plans.
  • Evaluate and adopt modern storage approaches (NVMeoF, RDMA, high-speed interconnects, cloud storage).
  • Guide incident response and root cause analysis, implementing durable preventive changes.
  • Partner with engineering, DevOps, and AI/ML teams to improve data pipelines and workflow performance.

Skills

People management
Storage systems
Incident response
Automation
Monitoring
Capacity planning
SRE principles
Storage protocols

Education

BS or MS in Computer Science, Storage Systems

Tools

Kubernetes
Terraform
Ansible
Puppet
Prometheus
InfluxDB
Elastic stack

Job description

NVIDIA is seeking a leader for Storage Production Engineering to design, deploy, and operate large-scale storage systems that power AI and HPC workloads. You will build a collaborative, learning-driven team and guide data pipelines and storage technologies across the platform.

You will own capacity planning, high availability, and incident response while partnering with engineering, DevOps, and AI/ML teams to improve performance and reliability.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Storage Production Engineering Lead - Scale & Reliability
Storage Production Engineering Lead - Scale & Reliability

NVIDIA • California (MO)

On-site
USD 272,000 - 431,250
Equity
Comprehensive benefits
Paid time off
Senior Storage Production Engineer: Low-Latency AI Cloud
Senior Storage Production Engineer: Low-Latency AI Cloud

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 176,000 - 276,000
Equity opportunities
Comprehensive benefits package
Senior Manager, Storage Production Engineering
Senior Manager, Storage Production Engineering

NVIDIA Corporation • Santa Clara (CA)

Hybrid
USD 272,000 - 431,000
Senior Manager, Storage Production Engineering
Senior Manager, Storage Production Engineering

NVIDIA • California (MO)

On-site
USD 272,000 - 432,000
Equity
Comprehensive benefits
Paid time off
Petabyte-Scale Storage Deployment Leader
Petabyte-Scale Storage Deployment Leader

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 248,000 - 397,000
Petabyte-Scale Storage Deployment Lead
Petabyte-Scale Storage Deployment Lead

NVIDIA • United States

On-site
USD 180,000 - 220,000
Senior Manager, Storage Engineering
Senior Manager, Storage Engineering

NVIDIA • United States

On-site
USD 180,000 - 220,000
Senior Storage Platform Engineer: Self-Service & IaC
Senior Storage Platform Engineer: Self-Service & IaC

NVIDIA Gruppe • Santa Clara (CA)

Hybrid
USD 168,000 - 334,000
Equity
Benefits
Lead Storage Software Engineer for AI GPU Clusters
Lead Storage Software Engineer for AI GPU Clusters

NVIDIA • Santa Clara (CA)

On-site
USD 224,000 - 431,000
Equity
Benefits
Storage Software Lead for AI GPU Clusters
Storage Software Lead for AI GPU Clusters

NVIDIA Corporation • Santa Clara (CA)

Hybrid
USD 224,000 - 431,000
Equity
Benefits