Storage Engineer (AI Infrastructure) - Hosting

Hamilton Barnes

San Francisco (CA)

Remote

USD 170,000 - 230,000

Full time

3 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Stock options
Remote work allowance

Job summary

Hamilton Barnes is seeking a Senior Storage Engineer to own the high-performance storage layer for large-scale GPU clusters and AI workloads. You will design, deploy, and operate storage platforms, collaborating with infrastructure, compute, and networking teams to scale performance.

The role focuses on optimizing distributed storage, troubleshooting bottlenecks, and driving resiliency and data protection through automation and observability in petabyte-scale environments.

Qualifications

  • Experience in storage engineering for HPC/AI infra or hyperscale data centers.
  • Hands-on expertise with VAST Data storage platforms preferred.
  • Strong knowledge of high-performance distributed storage architectures and parallel file systems.
  • Experience supporting GPU-intensive AI/ML workloads and high-throughput data environments.
  • Proficient Linux system administration.
  • Experience troubleshooting performance across storage, networking and compute.
  • Familiarity with NFS, RDMA, NVMe-oF, InfiniBand and modern storage networks.
  • Automation using Python, Bash, or Ansible.
  • Focus on scalability, resiliency and data protection in enterprise storage environments.

Responsibilities

  • Design, deploy, and operate large-scale high-performance storage platforms supporting AI and HPC workloads.
  • Manage and optimize distributed storage environments for large GPU training and inference clusters.
  • Work closely with platform, compute, and networking teams to ensure end-to-end infrastructure performance.
  • Troubleshoot storage bottlenecks, latency issues, throughput constraints, and data flow inefficiencies.
  • Contribute to storage architecture strategy, scalability planning, and operational best practices.
  • Automate storage provisioning, monitoring, and lifecycle management processes.
  • Support performance tuning across parallel file systems, object storage, and AI data pipelines.
  • Implement observability and capacity planning solutions for petabyte-scale environments.

Skills

Storage engineering
HPC/AI infra
VAST Data
Distributed storage
GPU AI workloads
Linux administration
NFS/RDMA/NVMe-oF
Python/Bash/Ansible
Scalability & resiliency

Tools

VAST Data
InfiniBand

Job description

Looking for a role with plenty of growth opportunities?

Join a rapidly scaling AI cloud infrastructure provider building next-generation GPU platforms for large-scale AI training, experimentation, and inference, significantly expanding operations across the United States alongside continued international growth.

This is a great opportunity for a Senior Storage Engineer to take ownership of the high-performance storage layer underpinning large GPU clusters and AI workloads at scale. The ideal candidate will work closely with infrastructure, networking, and platform engineering teams to design and optimize storage environments capable of supporting massive throughput, low latency, and highly parallelized AI training workloads.

Responsibilities:
  • Design, deploy, and operate large-scale high-performance storage platforms supporting AI and HPC workloads
  • Manage and optimize distributed storage environments for large GPU training and inference clusters
  • Work closely with platform, compute, and networking teams to ensure end-to-end infrastructure performance
  • Troubleshoot storage bottlenecks, latency issues, throughput constraints, and data flow inefficiencies
  • Contribute to storage architecture strategy, scalability planning, and operational best practices
  • Automate storage provisioning, monitoring, and lifecycle management processes
  • Support performance tuning across parallel file systems, object storage, and AI data pipelines
  • Implement observability and capacity planning solutions for petabyte-scale environments
Skills/Must Have:
  • Deep experience in storage engineering within HPC, AI infrastructure, hyperscale, or large-scale data center environments
  • Deep hands-on expertise with VAST Data storage platforms strongly preferred
  • Strong understanding of high-performance distributed storage architectures and parallel file systems
  • Experience supporting GPU-intensive AI/ML workloads and high-throughput data environments
  • Strong Linux systems administration skills
  • Experience troubleshooting performance across storage, networking, and compute layers
  • Familiarity with NFS, RDMA, NVMe-oF, InfiniBand, and modern storage networking concepts
  • Automation and scripting experience using Python, Bash, or Ansible preferred
  • Strong understanding of scalability, resiliency, and data protection in enterprise storage environments
Benefits:
  • Stock options
  • Remote working options and allowance
Salary:
  • Circa $200,000 base salary
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Member of Technical Staff - Storage Infrastructure
Member of Technical Staff - Storage Infrastructure

AI Chopping Block • San Francisco (CA), Northern (KY)

On-site
USD 150,000 - 300,000
Senior Storage Engineer
Senior Storage Engineer

Hydra Host • Miami (FL)

On-site
USD 120,000 - 160,000
Member of Technical Staff - Storage Infrastructure
Member of Technical Staff - Storage Infrastructure

Prime Intellect AI • San Francisco (CA)

On-site
USD 150,000 - 300,000
Member of Technical Staff - Storage Infrastructure
Member of Technical Staff - Storage Infrastructure

Prime Intellect • San Francisco (CA)

On-site
USD 150,000 - 300,000
Head of AI Data Center Infrastructure Platforms and Software
Head of AI Data Center Infrastructure Platforms and Software

Summit Group Solutions, LLC • United States

On-site
USD 150,000 - 350,000
Member of Technical Staff - Storage Infrastructure
Member of Technical Staff - Storage Infrastructure

Primeintellect • San Francisco (CA)

On-site
USD 150,000 - 300,000
Member of Technical Staff - Storage Infrastructure at Prime Intellect
Member of Technical Staff - Storage Infrastructure at Prime Intellect

Matcha • Northern (KY)

Hybrid
USD 150,000 - 300,000
Senior Solutions Engineer, AI Infrastructure
Senior Solutions Engineer, AI Infrastructure

VAST Data • New York (NY)

On-site
USD 150,000 - 200,000
Senior Storage Software Engineer, DGXC Data Services
Senior Storage Software Engineer, DGXC Data Services

NVIDIA • Town of Texas (WI)

On-site
USD 152,000 - 287,500
Equity
Benefits
Senior Storage Software Engineer, DGXC Data Services
Senior Storage Software Engineer, DGXC Data Services

NVIDIA • California (MO)

On-site
USD 184,000 - 287,500
Equity
Benefits