Storage Engineer - Hosting

Hamilton Barnes Associates Limited

United States

On-site

USD 170,000 - 230,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Stock options
Bonus 10%

Job summary

Hamilton Barnes Associates Limited is seeking a Storage Engineer to design and deploy high-performance AI storage systems for GPU-heavy workloads. You will architect parallel file systems, optimise data pipelines, and coordinate across global data centers to ensure 99.99% uptime.

The role requires deep Linux storage knowledge, GPU-stack awareness, and proficiency with InfiniBand networking. Responsibilities include automating storage provisioning with Terraform/Ansible and leading Tier-3

Qualifications

  • 5+ years of experience with high-performance storage solutions in a Linux-heavy environment.
  • Deep understanding of storage interaction with NVIDIA GPU stacks and ML training I/O patterns.
  • Hands-on experience with InfiniBand, RoCEv2, and NVMe-over-Fabrics.

Responsibilities

  • Design & Deploy AI Storage: Architect and implement high-performance parallel file systems optimized for GPU-heavy workloads.
  • Optimise Data Pipelines: Fine-tune storage performance for maximum GPUDirect Storage efficiency and low latency.
  • Manage Scale & Reliability: Build petabyte-scale storage clusters across multiple data centers with high uptime.
  • Infrastructure Integration: Collaborate with Network and Data Center teams to configure high-speed storage networking.
  • Automate Storage Ops: Develop Terraform, Ansible, or Python scripts to provision, monitor, and scale storage resources.
  • Troubleshoot I/O: Lead storage-related performance investigations across filesystem, network, or kernel.

Skills

Storage expertise
AI infrastructure knowledge
Networking proficiency
Systems automation
Linux internals

Tools

Terraform
Ansible
Python

Job description

Join a Founders Fund-backed NVIDIA cloud partner building the high-performance infrastructure that powers the world’s most ambitious AI research. In the world of GPUaaS, the bottleneck is rarely the compute; it’s the data.

You will be a Storage Engineer who understands that AI at scale requires more than just capacity; it requires massive throughput, ultra-low latency, and the ability to feed thousands of GPUs without a hiccup. You will architect and build the data layer that supports foundation model training and enterprise-grade production inference.

Responsibilities:
  • Design & Deploy AI Storage: Architect and implement high-performance parallel file systems (Weka, Lustre, or similar) optimised specifically for GPU-heavy workloads and multi-node training.
  • Optimise Data Pipelines: Fine-tune storage performance to ensure maximum GPUDirect Storage (GDS) efficiency, minimising latency between the storage fabric and the GPU memory.
  • Manage Scale & Reliability: Build and maintain petabyte-scale storage clusters across multiple global data centers, ensuring 99.99% uptime for mission-critical AI research labs.
  • Infrastructure Integration: Partner with Network and Data Center engineers to configure high-speed storage networking (InfiniBand/400G Ethernet) and ensure seamless backend connectivity.
  • Automate Storage Ops: Develop Terraform providers, Ansible playbooks, or Python scripts to automate the provisioning, monitoring, and scaling of storage resources.
  • Troubleshoot Complex I/O: Act as the Tier-3 lead for storage-related performance degradation, identifying root causes in the filesystem, network, or Linux kernel.
Skills/Must have:
  • Specialised Storage Expertise: 5+ years of experience with high-performance storage solutions (WekaIO, VAST Data, BeeGFS, or DDN) in a Linux-heavy environment.
  • AI Infrastructure Knowledge: Deep understanding of how storage interacts with NVIDIA GPU stacks (HGX/DGX) and the specific I/O patterns of ML training (checkpoints, small file reads, etc.).
  • Networking Proficiency: Hands-on experience with InfiniBand, RoCEv2, and NVMe-over-Fabrics (NVMe-oF).
  • Systems Automation: Strong scripting skills in Python, Go, or Bash, and experience with IaC tools like Terraform or Pulumi.
  • Linux Internals: Deep knowledge of the Linux storage stack, including XFS/ZFS, LVM, and kernel tuning for high-throughput networking.
Benefits:
  • 10% bonus
  • Stock options
Salary:
  • $200,000 base salary
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Storage Software Engineer, DGXC Data Services
Senior Storage Software Engineer, DGXC Data Services

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 152,000 - 288,000
Equity compensation
Benefits
Senior HPC Storage Engineer
Senior HPC Storage Engineer

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 184,000 - 288,000
Senior Storage Software Engineer - DGX Cloud
Senior Storage Software Engineer - DGX Cloud

2100 NVIDIA USA • Santa Clara (CA)

On-site
USD 224,000 - 432,000
Equity
Benefits
Senior Storage Software Engineer - DGX Cloud
Senior Storage Software Engineer - DGX Cloud

Nvidia Corporation • Santa Clara (CA)

On-site
USD 224,000 - 432,000
Equity
Benefits
Senior Manager, Storage Production Engineering
Senior Manager, Storage Production Engineering

NVIDIA • California (MO)

On-site
USD 272,000 - 432,000
Equity
Comprehensive benefits
Paid time off
Senior Storage Software Engineer, DGXC Data Services
Senior Storage Software Engineer, DGXC Data Services

NVIDIA • Town of Texas (WI)

On-site
USD 152,000 - 288,000
Equity
Benefits
Senior Storage Software Engineer, DGXC Data Services
Senior Storage Software Engineer, DGXC Data Services

NVIDIA • California (MO)

On-site
USD 184,000 - 288,000
Equity
Benefits
Senior Storage Software Engineer, DGXC Data Services
Senior Storage Software Engineer, DGXC Data Services

NVIDIA • United States

On-site
USD 152,000 - 288,000
Equity
Benefits
Senior Storage Software Engineer, DGXC Data Services
Senior Storage Software Engineer, DGXC Data Services

NVIDIA • North Carolina

On-site
USD 152,000 - 242,000
Equity compensation
Benefits
Senior Product Architect, Storage
Senior Product Architect, Storage

Visa Hunt • United States

On-site
USD 224,000 - 357,000
Equity
Benefits