Senior Storage Systems Architect for AI/ML & HPC

NVIDIA

Sydney

On-site

AUD 180,000 - 240,000

Full time

7 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

NVIDIA Storage Production Engineers design, implement, and support large-scale storage clusters, ensuring scalability and data integrity for GPU cloud services. You’ll develop monitoring, logging, and alerting systems and work with AI/ML workloads to optimize low-latency access and high throughput.

You will drive lifecycle improvements from design to deployment, maintain production infrastructure, and participate in on-call rotation while advancing security and efficiency in a fast-paced

Qualifications

  • BS degree or equivalent experience in Computer Science, Storage Systems, or a related technical field with 8+ years of practical experience.
  • Experience with distributed and high-performance storage solutions, including clustered and parallel file systems, distributed object storage, and enterprise-grade storage systems.
  • Solid understanding of block, file, and object storage technologies, including their scalability, reliability, and performance characteristics.
  • Experience with storage networking protocols such as NFS, SMB, iSCSI, S3, Fibre Channel, RDMA, and NVMe over Fabrics.
  • Expertise in algorithms, data structures, complexity analysis, software design, and automating maintenance of large-scale Linux-based storage systems.
  • Experience in one or more of the following: C/C++, Java, Python, Go, NodeJS, and Bash for storage automation, monitoring, and performance tuning.
  • Hands-on experience with infrastructure configuration management tools and observability stacks.

Responsibilities

  • Design, implement, and support large-scale storage clusters, ensuring scalability, high availability, and data integrity.
  • Develop and maintain storage monitoring, logging, and alerting systems to ensure proactive detection and resolution of performance issues.
  • Work with AI/ML workloads to improve storage architectures for low-latency access, efficient caching, and high-throughput performance.
  • Improve the lifecycle of storage services – from inception and design to deployment, operation, and continuous optimization.
  • Maintain production storage infrastructure by supervising availability, latency, and system health, leveraging predictive analytics and AI-driven automation.
  • Optimize storage efficiency through compression, deduplication, tiering strategies, and intelligent workload placement.
  • Scale storage systems sustainably using AI/ML-driven automation, policy-based tiering, and dynamic data migration techniques.
  • Ensure data security and compliance by implementing encryption, access controls, and auditing mechanisms for storage systems.
  • Practice sustainable incident response and blameless root cause analysis. Be part of an on-call rotation to support storage and production systems.

Skills

Distributed storage systems
High-performance storage
Linux
Automation
C/C++
Python
Go
CI/CD

Education

BS degree in Computer Science or related field

Tools

Ansible
Chef
Puppet
Terraform
InfluxDB/Prometheus/Grafana/Elastic

Job description

NVIDIA Storage Production Engineers design, implement, and support large-scale storage clusters, ensuring scalability and data integrity for GPU cloud services. You’ll develop monitoring, logging, and alerting systems and work with AI/ML workloads to optimize low-latency access and high throughput.

You will drive lifecycle improvements from design to deployment, maintain production infrastructure, and participate in on-call rotation while advancing security and efficiency in a fast-paced

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Storage Infrastructure Engineer for AI/ML & HPC
Senior Storage Infrastructure Engineer for AI/ML & HPC

Nvidia • New South Wales

On-site
AUD 180,000 - 240,000
Senior Storage Production Engineer - DGX Cloud
Senior Storage Production Engineer - DGX Cloud

NVIDIA • Sydney

On-site
AUD 180,000 - 240,000
Senior High-Performance Storage Architect for AI Infrastructure
Senior High-Performance Storage Architect for AI Infrastructure

NVIDIA • Sydney

On-site
AUD 251,000 - 363,000
Senior Solution Architect, AI Compute Engineer - NVIS
Senior Solution Architect, AI Compute Engineer - NVIS

NVIDIA • Sydney

On-site
AUD 120,000 - 160,000
AI Domain Architect (AI Storage)
AI Domain Architect (AI Storage)

World Wide Technology • City of Melbourne

On-site
AUD 180,000 - 240,000
AI Storage Architect: High-Performance NVMe & GPU AI
AI Storage Architect: High-Performance NVMe & GPU AI

World Wide Technology • Victoria

On-site
AUD 180,000 - 240,000
AI Storage Domain Architect – High-Performance NVMe & GPU AI
AI Storage Domain Architect – High-Performance NVMe & GPU AI

World Wide Technology • City of Melbourne

On-site
AUD 180,000 - 240,000
Senior Kubernetes Architect for AI Infrastructure
Senior Kubernetes Architect for AI Infrastructure

United States Digital Space LLC • Sydney

On-site
AUD 220,000 - 280,000
Senior HPC Network Engineer: GPU Scale & Fabric Design
Senior HPC Network Engineer: GPU Scale & Fabric Design

Pathway Search • Sydney

On-site
AUD 120,000 - 190,000
Senior Networking Solutions Architect - AI/HPC Cloud
Senior Networking Solutions Architect - AI/HPC Cloud

NVIDIA • Sydney

On-site
AUD 120,000 - 160,000