Senior Storage & Data Protection Engineer for HPC & AI

AHEAD

New York (NY)

On-site

USD 120,000 - 180,000

Full time

9 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

AHEAD is seeking a High-Performance Computing Storage Engineer to maintain storage technologies across managed services environments. You will support incident, problem, and change management, and work with teams to optimize, automate, and document storage infrastructure.

Responsibilities include administering Lustre/GPFS-like filesystems, tuning for AI workloads, and participating in on-call rotations, with emphasis on customer communication and cross-functional collaboration.

Qualifications

  • 5+ years of HPC, AI infrastructure, or large-scale storage engineering.
  • Bachelor's degree or equivalent Information Systems or related field.
  • Strong Linux systems administration experience.
  • Hands-on experience configuring, managing, and tuning distributed or parallel filesystems.
  • Experience tuning storage for performance-sensitive workloads.
  • Knowledge of HPC schedulers such as Slurm and/or container platforms such as Kubernetes.
  • Familiarity with high-speed interconnects such as InfiniBand or RDMA.
  • Ability to troubleshoot complex issues across storage, compute, and networking layers.
  • Experience with machine learning or data science workflows in HPC environments.
  • Managed Services or consulting experience.
  • Strong background with customer service.
  • High level problem-solving and communication skills.
  • Strong oral and written communications skills.

Responsibilities

  • Provide enterprise-level operational support to Managed Services customers for incident, problem, and change management activities.
  • Administer parallel and distributed filesystems such as Lustre, GPFS, BeeGFS, Ceph, Weka, or Vast.
  • Optimize storage performance, throughput, metadata operations, and data locality for AI training and inference.
  • Build and maintain automation for storage provisioning, monitoring, alerting, quota management, and lifecycle operations.
  • Plan and perform maintenance activities.
  • Assess customer environments for performance and design issues and propose resolutions.
  • Work across technical teams to troubleshoot complex infrastructure issues.
  • Create and maintain detailed documentation.
  • Serve as a subject matter expert and escalation point for storage technologies.
  • Work with vendors to resolve storage issues.
  • Communicate with customers and internal team with transparency.
  • Support data movement workflows including ingest, replication, caching, tiering, and archiving.
  • Troubleshoot storage, Linux, network, and I/O bottlenecks across storage clusters and fabrics.
  • Partner with infrastructure, platform, and research teams to support production AI/HPC workloads.
  • Evaluate new storage architectures and technologies for scalability, resilience, and cost efficiency.
  • Participate in on-call rotation.

Skills

HPC infrastructure
Linux administration
Troubleshooting
Customer service
On-call support
AI workflows

Education

Bachelor's degree in Information Systems

Tools

Slurm
Kubernetes
InfiniBand
RDMA
Ceph
GPFS
Terraform
Ansible
Helm
Prometheus
Grafana

Job description

AHEAD is seeking a High-Performance Computing Storage Engineer to maintain storage technologies across managed services environments. You will support incident, problem, and change management, and work with teams to optimize, automate, and document storage infrastructure.

Responsibilities include administering Lustre/GPFS-like filesystems, tuning for AI workloads, and participating in on-call rotations, with emphasis on customer communication and cross-functional collaboration.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Engineer, Storage and Data Protection
Engineer, Storage and Data Protection

AHEAD • New York (NY)

On-site
USD 120,000 - 180,000
Senior AI-Driven HPC Storage Architect
Senior AI-Driven HPC Storage Architect

NVIDIA Corporation • Santa Clara (CA), Northern (KY)

Hybrid
USD 148,000 - 288,000
Equity
Comprehensive benefits
Senior Storage & Cloud Infrastructure Consultant
Senior Storage & Cloud Infrastructure Consultant

AHEAD • Nashville (TN)

On-site
USD 110,000 - 170,000
Medical Insurance
401(k)
Paid time off
+1
HPC Storage Architect for Large-Scale Data & AI
HPC Storage Architect for Large-Scale Data & AI

UCSF Health • San Francisco (CA)

On-site
USD 140,000 - 190,000
Senior AI Storage Engineer — High-Performance GPU Data Fabric
Senior AI Storage Engineer — High-Performance GPU Data Fabric

Hamilton Barnes Associates Limited • United States

On-site
USD 170,000 - 230,000
Stock options
Bonus 10%
AI Factory Storage Architect for HPC-Scale
AI Factory Storage Architect for HPC-Scale

NVIDIA Corporation • California (MO)

Hybrid
USD 148,000 - 288,000
Senior AI/HPC Solutions Architect – Lustre Storage
Senior AI/HPC Solutions Architect – Lustre Storage

NetApp, Inc. • United States

On-site
USD 197,000 - 256,000
Health Insurance
Stock options
Paid time off
+1
Senior HPC Storage SRE — On‑Prem/Cloud, Equity Options
Senior HPC Storage SRE — On‑Prem/Cloud, Equity Options

Nvidia Corporation in • Santa Clara (CA)

On-site
USD 168,000 - 322,000
Senior HPC Storage SRE: On-Prem + Cloud Automation
Senior HPC Storage SRE: On-Prem + Cloud Automation

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 168,000 - 322,000
AI Storage Architect for High-Performance GPU Cloud
AI Storage Architect for High-Performance GPU Cloud

Bitdeer (NASDAQ: BTDR) • Austin (TX)

On-site
USD 140,000 - 200,000