Senior HPC Storage Engineer

The Regents of the University of California on behalf of their Los Angeles Campus

Los Angeles (CA)

Hybrid

USD 150,000 - 230,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

UCLA’s Office of Advanced Research Computing (OARC) seeks a Senior HPC Storage Engineer to lead design, deployment, and operation of large-scale storage systems supporting data-intensive research, AI/ML, and multi-institution collaborations. You will evaluate storage tech, deploy petabyte-scale platforms, and ensure secure, reliable services across campus, cloud, and federated HPC environments.

Join a collaborative team advancing national-scale research data infrastructure through the iDLab

Qualifications

  • 7 years or more experience in storage environments with responsibility for large-scale production storage services (Required).
  • Experience with petabyte-scale storage, federated storage, multi-site research infrastructure, or HPC/cloud-integrated storage (Preferred).
  • Advanced knowledge of HPC, data science, and cyberinfrastructure supporting large-scale workloads (Required).
  • Knowledge of scale-out, parallel, distributed, object, federated storage architectures and performance considerations (Required).
  • Hands-on experience with large-scale storage platforms (Lustre, VAST Data, GPFS/Spectrum Scale, Ceph, BeeGFS, MinIO) (Required).
  • Experience operating large-scale production storage systems (monitoring, capacity planning, upgrades, tuning, incident response, backup, DR) (Required).
  • Advanced Linux administration including kernel troubleshooting and tuning (Required).
  • Scripting/automation using Bash/Python; familiarity with Ansible and Git (Required).
  • Experience integrating storage with identity providers (AD/LDAP/OIDC/SAML/Globus) (Required).
  • Ability to design and execute benchmarks and document results (Required).
  • Ability to design/support shared research storage services and data movement (Preferred).
  • Strong communication with technical and leadership audiences (Required).
  • Ability to lead technical teams and manage multi-month projects (Required).

Responsibilities

  • Lead design, deployment, and operation of petabyte-scale storage systems for data-intensive research.
  • Evaluate emerging storage technologies and architectures for campus, cloud, and federated HPC environments.
  • Develop secure, reliable storage services and perform performance optimization.
  • Conduct storage benchmarks, validate vendor claims, and document results for engineering/executive audiences.
  • Coordinate with researchers and partner institutions; contribute to national-scale data infrastructure initiatives.
  • Support deployment and upgrades, incident response, backup, replication, and disaster recovery for storage platforms.

Skills

Large-scale storage
HPC storage
Storage architectures
Lustre
GPFS/Spectrum Scale
Ceph
BeeGFS
MinIO
Performance tuning
Linux administration
Scripting (Bash/Python)
Identity integration (AD/LDAP/OIDC/SAM

Education

Bachelor's Degree in Computer Science/related field
Master's Degree (Preferred)
PhD (Preferred)

Tools

Ansible
Git
Python

Job description

Special Instructions to Applicants

This position will consider both hybrid and fully remote candidates. On-site presence may be required as needed based on operational, maintenance, or project requirements. Participation in an on-call rotation for emergency hardware and software support is required. Occasional evenings and weekends may be required to support maintenance windows or respond to incidents. Required to accommodate scheduled based on Pacific Standard Time.

Department Summary

Advanced Research Computing melds expert staff and technical infrastructure to amplify and accelerate the impact of UCLA research in the age of networked data and computation. OARC’s expertise and resources are available to all UCLA researchers engaged in digital research and scholarship. We work with faculty, student, and postdoctoral researchers; instructors; and staff and administrators. OARC is a relationship-building organization. We enable digital scholarship through collaborations, partnerships, and networked communities to advance cutting-edge research capabilities at UCLA and beyond. OARC supports and enhances the university mission of education, research, and service through the development and execution of innovative and sustainable technology practices, programs, services, infrastructure, policies, and partnerships.

Position Summary

Join UCLA’s Office of Advanced Research Computing (OARC) and help build the research data infrastructure that powers scientific discovery at UCLA and across a national network of high-performance computing centers. We are seeking a Senior HPC Storage Engineer to lead the design, deployment, and operation of large-scale storage systems supporting data-intensive research, AI/ML, interactive computing, and multi-institutional collaborations.

This senior technical role combines architecture with hands-on engineering. You will evaluate emerging storage technologies, lead deployment of petabyte-scale storage platforms, optimize performance, and develop secure, reliable storage services spanning campus, cloud, and federated HPC environments. As part of UCLA’s NSF-supported iDLab initiative, you will help build a national-scale research data infrastructure that enables seamless access to data across geographically distributed computing resources.

At OARC, you will work alongside leading researchers, engineers, and partner institutions while influencing the long-term direction of research cyberinfrastructure at UCLA. This is an opportunity to solve challenging technical problems, work with advanced storage technologies, and make a lasting impact on research computing at both the campus and national levels.

Salary & Compensation

*UCLA provides a full pay range. Actual salary offers consider factors, including budget, prior experience, skills, knowledge, abilities, education, licensure and certifications, and other business considerations. Salary offers at the top of the range are not common. Visit UC Benefit package to discover benefits that start on day one, and UC Total Compensation Estimator to calculate the total compensation value with benefits.

Qualifications
  • 7 years or more Experience in research, enterprise, or hyperscale storage environments with responsibility for large-scale production storage services. (Required)
  • Large scale storage Experience with petabyte-scale storage, federated storage, multi-site research infrastructure, or HPC/cloud-integrated storage services. (Preferred)
  • 1. Advanced knowledge of high-performance computing, data science, and cyberinfrastructure environments supporting large-scale research workloads (Required)
  • 2. Demonstrated knowledge of scale-out, parallel, distributed, object, and federated storage architectures, including performance characteristics, consistency models, caching strategies, and operational tradeoffs. (Required)
  • 3. Hands-on experience architecting, deploying, operating, or substantially supporting one or more large-scale storage platforms such as Lustre, VAST Data, GPFS/Spectrum Scale, Ceph, BeeGFS, WekaFS, or MinIO. Experience with multiple platforms and petabyte-scale deployments is preferred. (Required)
  • 4. Demonstrated experience operating large-scale production storage systems, including monitoring, capacity planning, lifecycle management, upgrades, performance tuning, incident response, backup, replication, and disaster recovery. (Required)
  • 5. Advanced Linux systems administration skills, including kernel-level troubleshooting, performance profiling, storage hardware diagnostics, and tuning across InfiniBand, RoCE, and Ethernet fabrics. (Required)
  • 6. Demonstrated experience diagnosing storage performance problems across Linux clients, metadata services, storage servers, object services, network fabrics, protocol layers, and application I/O patterns. (Required)
  • 7. Proficiency in scripting and automation using Bash, Python, or similar languages; familiarity with configuration management tools such as Ansible and version control such as Git. (Required)
  • 8. Experience integrating storage systems with local, campus-wide, and federated identity providers such as Active Directory, LDAP, CILogon, OIDC, SAML, or Globus Auth. (Required)
  • 9. Demonstrated ability to design and execute storage benchmarks, validate vendor claims, characterize representative scientific workloads, and document results for engineering and executive audiences. (Required)
  • 10. Demonstrated ability to design, operate, or support shared research storage services, including allocation models, user-facing service delivery, large-scale data movement, and escalation support for complex research workflows. (Preferred)
  • 11. Demonstrated ability to communicate complex technical information clearly to technical staff, researchers, leadership, vendors, and external research and education audiences. (Required)
  • 12. Demonstrated ability to work independently and collaboratively, manage competing priorities, lead technical working groups, and sustain production service quality while delivering multi-month projects. (Required)
Education, Licenses, Certifications & Personal Affiliations
  • Bachelor's Degree Bachelor's degree in Computer Science, Computational Science, Data Science, Engineering, or a related field; or equivalent combination of education and experience. (Required)
  • Master's Degree Master's in Computer Science, Computational Science, Data Science, Engineering, or a related field; or equivalent advanced professional experience. (Preferred) Or
  • PhD PhD in Computer Science, Computational Science, Data Science, Engineering, or a related field; or equivalent advanced professional experience. (Preferred)
Special Conditions for Employment
  • Background Check: Continued employment is contingent upon the completion of a satisfactory background investigation.
  • Live Scan Background Check: A Live Scan background check must be completed prior to the start of employment.
  • This position will consider both hybrid and fully remote candidates. On-site presence may be required as needed based on operational, maintenance, or project requirements. Participation in an on-call rotation for emergency hardware and software support is required. Occasional evenings and weekends may be required to support maintenance windows or respond to incidents. (Required)
Schedule

8:00 a.m. to 5:00 p.m

Union/Policy Covered

RP-Research and Public Service PR

Complete Position Description

https://universityofcalifornia.marketpayjobs.com/ShowJob.aspx?EntityID=38&JDName=Computational%20and%20Data%20Science%20Research%20Specialist%204%20RP%20(TBD_1000126)

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior HPC Storage Engineer
Senior HPC Storage Engineer

University of California Los Angeles • Los Angeles (CA)

Hybrid
USD 150,000 - 210,000
Senior HPC Storage Engineer
Senior HPC Storage Engineer

University of California - Los Angeles (UCLA) • Los Angeles (CA)

Hybrid
USD 150,000 - 230,000
Senior HPC Storage Engineer
Senior HPC Storage Engineer

UCLA • Los Angeles (CA)

Hybrid
USD 150,000 - 200,000
Senior HPC Storage Architect - Remote & On-Call
Senior HPC Storage Architect - Remote & On-Call

The Regents of the University of California on behalf of their Los Angeles Campus • Los Angeles (CA)

Hybrid
USD 150,000 - 230,000
Senior HPC Storage Architect — Remote/Hybrid
Senior HPC Storage Architect — Remote/Hybrid

UCLA • Los Angeles (CA)

Hybrid
USD 150,000 - 200,000
Senior HPC Storage Architect - Remote/Hybrid
Senior HPC Storage Architect - Remote/Hybrid

University of California Los Angeles • Los Angeles (CA)

Hybrid
USD 150,000 - 210,000
Research Computing Storage Administrator
Research Computing Storage Administrator

CU Boulder • Boulder (CO)

Hybrid
USD 71,000 - 78,000
Medical insurance
Dental insurance
Retirement plans
+3
Research Computing Storage Administrator
Research Computing Storage Administrator

The Chronicle Of Higher Education, Inc. • Boulder (CO)

Hybrid
USD 71,000 - 78,000
Hybrid work schedule
Research Computing Storage Administrator
Research Computing Storage Administrator

University of Colorado Boulder • Boulder (CO)

Hybrid
USD 71,000 - 78,000
Medical, dental, retirement plans
Paid time off
Tuition assistance
+2
Research Storage Engineer Senior
Research Storage Engineer Senior

University of Michigan • Ann Arbor (MI)

Hybrid
USD 103,000 - 114,000
Generous time off policy
Two-for-one retirement matching
Comprehensive health insurance
+5