Storage Production Engineering Lead – Remote

NVIDIA Corporation

Santa Clara (CA)

On-site

USD 272,000 - 431,250

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity
401(k) plan
Employee stock purchase plan
Paid time off

Job summary

NVIDIA is seeking a Storage Production Engineering Leader to guide a team that designs and operates large-scale storage systems. You’ll own capacity planning, DR, and performance tuning while partnering with engineering, DevOps, and AI teams to optimize data pipelines and access patterns.

You will apply IaC, automation, and modern storage technologies to improve reliability and efficiency across NVIDIA’s storage platforms, including NVMe and cloud-based approaches, with a focus on high

Qualifications

  • BS or MS in Computer Science, Storage Systems, or related technical field, or equivalent experience.
  • 12+ years in large scale storage architecture, operations, or production engineering.
  • 6+ years of people management or technical leadership in storage/infrastructure.
  • Hands-on with parallel file systems (Lustre/GPFS), distributed storage (Ceph/MinIO), and enterprise object/NAS (S3, NetApp, Pure).
  • Knowledge of block, file, object storage performance tuning and high availability.
  • Experience with storage networking protocols (NFS, SMB, iSCSI, Fibre Channel, RDMA, NVMe-oF).
  • Automation and IaC with Terraform, Ansible, Puppet.
  • Monitoring/observability tools (Prometheus, InfluxDB, Elastic).

Responsibilities

  • Lead a team of Storage Production Engineers to design and operate large-scale storage systems.
  • Design, deploy, and improve distributed storage, parallel file systems, and object storage.
  • Use automation and monitoring to improve reliability and efficiency.
  • Own capacity planning, data lifecycle management, and DR planning for storage.
  • Guide incident response and root cause analysis to prevent repeats.
  • Collaborate with engineering, DevOps, and AI teams to optimize data pipelines and performance.

Skills

Leadership
Storage architecture
On-call / incident response
SRE/Production Engineering
Networking/storage protocols
Infrastructure as code
Monitoring/observability

Education

BS/MS in CS or related

Tools

Lustre
GPFS
Ceph
MinIO
S3 compatible storage
Terraform
Ansible
Puppet

Job description

NVIDIA is seeking a Storage Production Engineering Leader to guide a team that designs and operates large-scale storage systems. You’ll own capacity planning, DR, and performance tuning while partnering with engineering, DevOps, and AI teams to optimize data pipelines and access patterns.

You will apply IaC, automation, and modern storage technologies to improve reliability and efficiency across NVIDIA’s storage platforms, including NVMe and cloud-based approaches, with a focus on high

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Storage Production Engineering Lead - Scale & Reliability
Storage Production Engineering Lead - Scale & Reliability

NVIDIA • California (MO)

On-site
USD 272,000 - 432,000
Equity
Comprehensive benefits
Paid time off
Storage Production Engineering Lead for AI-Driven Systems
Storage Production Engineering Lead for AI-Driven Systems

NVIDIA AI • Santa Clara (CA)

On-site
USD 272,000 - 432,000
Equity
Wellness programs
Senior Manager, Storage Production Engineering
Senior Manager, Storage Production Engineering

NVIDIA • California (MO)

On-site
USD 272,000 - 432,000
Equity
Comprehensive benefits
Paid time off
Senior Storage Production Engineer: Low-Latency AI Cloud
Senior Storage Production Engineer: Low-Latency AI Cloud

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 176,000 - 276,000
Equity opportunities
Comprehensive benefits package
Senior Manager, Storage Production Engineering
Senior Manager, Storage Production Engineering

NVIDIA AI • Santa Clara (CA)

On-site
USD 272,000 - 432,000
Equity
Wellness programs
Director of Storage Acceleration & Platform Engineering
Director of Storage Acceleration & Platform Engineering

NVIDIA • Redmond (WA)

On-site
USD 320,000 - 489,000
Equity
Benefits
Senior Manager, Storage Production Engineering
Senior Manager, Storage Production Engineering

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 272,000 - 432,000
Equity
401(k) plan
Employee stock purchase plan
+1
Senior AI Storage Software Engineer — Scale GPU Clusters
Senior AI Storage Software Engineer — Scale GPU Clusters

Nvidia Corporation • Santa Clara (CA)

On-site
USD 224,000 - 432,000
Equity
Benefits
Senior AI Data Storage Engineer - Cloud-Native, Equity
Senior AI Data Storage Engineer - Cloud-Native, Equity

NVIDIA • North Carolina

On-site
USD 152,000 - 242,000
Equity compensation
Benefits
HPC Storage Architect for AI/GPU Infrastructure
HPC Storage Architect for AI/GPU Infrastructure

NVIDIA • Santa Clara (CA)

On-site
USD 184,000 - 287,500
Comprehensive benefits package
Equity options