Senior AI Cloud Storage Systems Engineer

NVIDIA Corporation

Santa Clara (CA)

On-site

USD 176,000 - 334,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

NVIDIA Corporation seeks a Production Storage Engineer to design, deploy, and optimize large-scale storage clusters for GPU-accelerated AI/ML workloads in Santa Clara, CA. You will build and maintain monitoring, logging, and alerting, ensuring data integrity, low latency, and high availability across distributed storage systems.

You will work with cutting-edge storage technologies, implement automation frameworks, and collaborate with software, systems, and hardware teams to enhance storage

Qualifications

  • BS degree in Computer Science, Storage Systems, or a related field with 8+ years of practical experience.
  • Experience with distributed and high-performance storage solutions, including clustered and parallel file systems, distributed object storage, and enterprise-grade storage systems.
  • Solid understanding of block, file, and object storage technologies and their scalability, reliability, and performance characteristics.
  • Experience with storage networking protocols such as NFS, SMB, iSCSI, S3, Fibre Channel, RDMA, and NVMe over Fabrics.
  • Expertise in algorithms, data structures, software design, and automating maintenance of large-scale Linux-based storage systems.
  • Experience in one or more of the following: C/C++, Java, Python, Go, NodeJS, and Bash for storage automation, monitoring, and performance tuning.
  • Hands-on experience with infrastructure configuration management tools like Ansible, Chef, Puppet, and Terraform for automating storage deployments.
  • Experience with observability and tracing tools like InfluxDB, Prometheus, Grafana, and the Elastic stack for monitoring storage system health.
  • Excellent written and oral communication, teamwork, and commitment to quality.

Responsibilities

  • Design, implement, and support large-scale storage clusters with high availability and data integrity.
  • Develop and maintain storage monitoring, logging, and alerting systems for proactive issue resolution.
  • Work with AI/ML workloads to improve storage architectures for low-latency access and high throughput.
  • Improve the lifecycle of storage services from design to deployment and optimization.
  • Support storage services with system build consulting, automation frameworks, capacity management, and launch reviews.
  • Maintain production storage infrastructure by monitoring availability, latency, and health with predictive analytics and AI-driven automation.
  • Optimize storage efficiency through compression, deduplication, tiering, and intelligent workload placement.
  • Scale storage systems using AI/ML-driven automation and dynamic data migration techniques.
  • Ensure data security with encryption, access controls, and auditing for storage systems.
  • Participate in on-call rotations for storage and production systems.

Skills

Distributed storage
C/C++
Python
Go
Bash
Linux
CI/CD
Terraform
Ansible
Kubernetes

Education

BS in CS or related

Tools

OpenStack
Kubernetes
Terraform
Ansible
Prometheus
Grafana
InfluxDB
ELK

Job description

NVIDIA Corporation seeks a Production Storage Engineer to design, deploy, and optimize large-scale storage clusters for GPU-accelerated AI/ML workloads in Santa Clara, CA. You will build and maintain monitoring, logging, and alerting, ensuring data integrity, low latency, and high availability across distributed storage systems.

You will work with cutting-edge storage technologies, implement automation frameworks, and collaborate with software, systems, and hardware teams to enhance storage

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Storage Production Engineer: Low-Latency AI Cloud
Senior Storage Production Engineer: Low-Latency AI Cloud

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 176,000 - 276,000
Equity opportunities
Comprehensive benefits package
Senior AI Data Storage Engineer - Cloud-Native, Equity
Senior AI Data Storage Engineer - Cloud-Native, Equity

NVIDIA • North Carolina

On-site
USD 152,000 - 241,500
Equity compensation
Benefits
Senior AI Storage Architect for High-Performance Systems
Senior AI Storage Architect for High-Performance Systems

NVIDIA • Massachusetts

On-site
USD 148,000 - 288,000
Equity
Benefits package
Remote: Head of Storage Production Engineering
Remote: Head of Storage Production Engineering

NVIDIA Corporation • Santa Clara (CA)

Hybrid
USD 272,000 - 431,000
Senior AI Storage Architect for High-Performance Systems
Senior AI Storage Architect for High-Performance Systems

NVIDIA • California (MO)

On-site
USD 184,000 - 288,000
Equity
Comprehensive benefits
Senior Storage Systems Engineer for AI & Cloud Data
Senior Storage Systems Engineer for AI & Cloud Data

NVIDIA • Town of Texas (WI)

On-site
USD 152,000 - 287,500
Equity
Benefits
Senior AI Factory Storage Engineer | HPC, NVMeoF & Cloud
Senior AI Factory Storage Engineer | HPC, NVMeoF & Cloud

NVIDIA Corporation • United States

Remote
USD 150,000 - 190,000
Senior AI Storage Architect — High-Performance Infra
Senior AI Storage Architect — High-Performance Infra

NVIDIA Corporation • Santa Clara (CA)

Remote
AUD 260,000 - 346,000
Storage Software Lead for AI GPU Clusters
Storage Software Lead for AI GPU Clusters

NVIDIA Corporation • Santa Clara (CA)

Hybrid
USD 224,000 - 431,000
Equity
Benefits
Lead High-Performance AI Storage Engineer
Lead High-Performance AI Storage Engineer

NVIDIA • United States

Remote
USD 140,000 - 180,000