Senior HPC Storage SRE — On-Prem + Cloud, Equity

NVIDIA

Santa Clara (CA)

On-site

USD 208,000 - 334,000

Full time

5 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Equity
Benefits package

Job summary

NVIDIA is seeking a Senior Site Reliability Engineer focused on HPC storage to design, implement, and optimize on-prem storage with cloud integration. You will craft distributed storage, build automation tooling, and ensure efficient operations of our growing IT ecosystem.

You will work with Lustre/GPFS, NetApp, Pure Storage, and Cloudian MinIO, leveraging AWS/Azure/GCP and monitoring stacks to keep systems reliable and scalable in a fast-paced HPC environment.

Qualifications

  • 8+ years of HPC storage experience or equivalent; strong performance tuning.
  • Experience with NetApp/Pure Storage and S3-based storage (Cloudian MinIO).
  • Proficient with Lustre or GPFS file systems.
  • Python/Bash/Golang scripting and automation.
  • Experience in AWS/Azure/GCP environments and monitoring stacks.

Responsibilities

  • Design and implement on-prem HPC storage with cloud integration.
  • Develop scalable storage solutions for data-intensive apps.
  • Build automation tooling for deployment and management.
  • Document procedures, evaluations, and best practices.
  • Collaborate with engineering to meet infrastructure needs.
  • Influence deployment methodologies for performance and cost efficiency.

Skills

HPC storage
Cloud computing
Automation tooling
Distributed file systems
Python/Bash/Golang
Cloud environments (AWS/Azure/GCP)
Monitoring stacks
Communication skills

Education

BS in Computer Science
MS or PhD a plus

Tools

NetApp
Pure Storage
Cloudian MinIO
Lustre
GPFS
Docker
Kubernetes
Slurm
PBS
LSF
Prometheus
Grafana
Elasticsearch
Kibana
Splunk
Zabbix

Job description

NVIDIA is seeking a Senior Site Reliability Engineer focused on HPC storage to design, implement, and optimize on-prem storage with cloud integration. You will craft distributed storage, build automation tooling, and ensure efficient operations of our growing IT ecosystem.

You will work with Lustre/GPFS, NetApp, Pure Storage, and Cloudian MinIO, leveraging AWS/Azure/GCP and monitoring stacks to keep systems reliable and scalable in a fast-paced HPC environment.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior HPC Storage SRE — On-Prem & Cloud, Equity
Senior HPC Storage SRE — On-Prem & Cloud, Equity

NVIDIA AI • Santa Clara (CA)

On-site
USD 168,000 - 334,000
Equity
Benefits package
Senior HPC Storage SRE — On‑Prem/Cloud, Equity Options
Senior HPC Storage SRE — On‑Prem/Cloud, Equity Options

Nvidia Corporation in • Santa Clara (CA)

On-site
USD 168,000 - 322,000
Senior SRE: Global HPC & Cloud Reliability
Senior SRE: Global HPC & Cloud Reliability

NVIDIA AI • Durham (NC)

On-site
USD 184,000 - 288,000
Equity
Benefits package
Senior HPC Storage Architect - Equity Eligible
Senior HPC Storage Architect - Equity Eligible

NVIDIA AI • Santa Clara (TX)

On-site
USD 184,000 - 357,000
Senior SRE: Global HPC & Multi-Cloud Reliability
Senior SRE: Global HPC & Multi-Cloud Reliability

NVIDIA Corporation • Durham (CA), Northern (KY)

Hybrid
USD 152,000 - 288,000
Senior Site Reliability Engineer - Storage
Senior Site Reliability Engineer - Storage

NVIDIA AI • Santa Clara (CA)

On-site
USD 168,000 - 334,000
Equity
Benefits package
Senior Storage Engineering Leader - Petabyte-Scale, Equity
Senior Storage Engineering Leader - Petabyte-Scale, Equity

Nvidia Corporation in • Santa Clara (CA)

On-site
USD 248,000 - 397,000
Senior Site Reliability Engineer - Storage
Senior Site Reliability Engineer - Storage

Nvidia Corporation in • Santa Clara (CA)

On-site
USD 168,000 - 322,000
Senior Site Reliability Engineer - Storage
Senior Site Reliability Engineer - Storage

NVIDIA • Santa Clara (CA)

On-site
USD 208,000 - 334,000
Equity
Benefits package
Senior Staff SRE: Global On-Prem & Cloud Infra Lead
Senior Staff SRE: Global On-Prem & Cloud Infra Lead

NVIDIA • United States

On-site
USD 200,000 - 322,000