Compensation: Competitive Base Salary + Performance Bonus
Overview
Our client is seeking a Senior HPC Storage Engineer to design, deploy, optimize, and support large-scale storage infrastructure powering advanced HPC, AI, and data-intensive workloads.
This role will join an HPC storage engineering team responsible for more than 200PB of high-performance storage, with a significant focus on VAST Data environments.
The engineer will help scale storage capacity, improve performance, automate infrastructure operations, and maintain highly available storage platforms supporting mission-critical compute environments.
Key Responsibilities
- Design, deploy, administer, and scale VAST Data and other high-performance storage environments.
- Support multi-petabyte storage infrastructure with a focus on availability, throughput, scalability, and low-latency access.
- Serve as a technical resource for storage architecture, configuration, upgrades, capacity planning, and troubleshooting.
- Support storage lifecycle activities including deployments, expansions, migrations, upgrades, and technology refreshes.
- Develop storage engineering standards and operational best practices.
Performance & Optimization
- Monitor and optimize storage performance across HPC and AI environments.
- Troubleshoot throughput, latency, capacity, metadata, and workload-related issues.
- Identify bottlenecks and recommend architectural or configuration improvements.
- Partner with compute, networking, and application teams to optimize end-to-end data workflows.
- Analyze utilization and workload behavior to improve infrastructure efficiency.
Automation
- Develop automation for storage provisioning, configuration, monitoring, validation, and lifecycle management.
- Build repeatable workflows using Python and modern automation practices.
- Support Infrastructure-as-Code and DevOps approaches across storage operations.
- Reduce manual processes through scripting, APIs, and automation tooling.
Operations & Troubleshooting
- Troubleshoot complex issues across storage, hardware, software, networking, and workload layers.
- Support highly available production storage environments.
- Perform root-cause analysis and implement long-term corrective actions.
- Participate in capacity planning, infrastructure expansion, and operational readiness.
Cross-Functional Engineering
- Work closely with HPC, platform, network, software, and infrastructure teams.
- Partner with researchers and application teams to understand workload and performance requirements.
- Coordinate with storage vendors on architecture, escalations, upgrades, and product issues.
- Work with technologies from VAST Data, Dell, HPE, Rubrik, and similar vendors.
Required Qualifications
- Strong hands-on experience designing, deploying, or supporting enterprise or HPC storage infrastructure.
- Experience with VAST Data strongly preferred.
- Experience with distributed, clustered, or scale-out storage platforms.
- Strong understanding of file, block, and object storage.
- Experience supporting multi-petabyte or rapidly scaling storage environments.
- Strong knowledge of storage performance, capacity planning, redundancy, availability, and data protection.
- Experience troubleshooting across storage, compute, and networking layers.
- Ability to support highly available production environments.
Technical Skills
- VAST Data administration, architecture, or engineering.
- Experience with Dell PowerScale / Isilon, HPE, or similar scale-out storage technologies.
- Python or comparable scripting experience.
- Experience with APIs, automation frameworks, or Infrastructure-as-Code.
- Understanding of high-speed networking as it relates to storage performance.
- Experience with monitoring and infrastructure management platforms.
Preferred Experience
- HPC, AI, machine learning, or research computing environments.
- Storage infrastructure at 100PB+ scale.
- Large GPU or CPU compute environments.
- High-performance networking technologies.
- Large-scale backup, replication, and data protection platforms.