Contribute to the strategic objectives of the High Performance Computing environment in the Advanced Computing Technologies department for our client, a leading business aviation aircraft manufacturer. Develop operational plans, goals, and strategies to best serve the Computational Fluid Dynamics, Simulation, and Modeling Engineering business units. Provide technical oversight and operational support to ensure continued sustained functionality of the HPC environment, and drive optimum integration of scientific applications to HPC technology.
Principal Duties and Responsibilities
- Own day-to-day operations of the production HPC cluster
- Support end users running applications on the HPC cluster
- Provide third-level support for engineering workstations and remote visualization systems
- Manage, maintain, monitor, and control interactive and batch processes (scheduled and unscheduled)
- Ensure batch processing and backups complete correctly and on schedule
- Improve processing capabilities and efficiency through system tuning and optimization
- Monitor utilization and report on tuning initiatives
- Perform proactive failure trend analysis and root cause analysis for system failures
- Produce trend reports and follow escalation procedures for production issues
- Monitor and adjust systems to support proper application execution
- Provide technical solutions meeting performance and processing objectives
- Perform upgrades per corporate policy and industry best practices
- Lead HPC Administrators during system upgrades and outages
- Create thorough upgrade plans
- Support introduction of new technologies to improve capability, productivity, and cost of ownership
- Participate in HPC technical solution design
- Evaluate efficiency, technology effectiveness, and interoperability on an ongoing basis
- Maintain relationships with hardware/software vendors
- Work multiple operational windows as required; provide 24x7 on-call support
- Support development and implementation of technical, hardware, and software standards
Experience and Education Requirements
- Bachelor's degree in Engineering, Computer Science, Information Technology, or related field required (or equivalent combination of education/experience)
- Master's degree may offset 1 year of experience; PhD may offset 2 years
- 7 years of experience in an HPC or scientific computing environment, including installation, configuration, and maintenance of RPM-based Linux distributions (RedHat, SuSE)
- Experience managing InfiniBand-based Linux HPC clusters, high-performance parallel storage, and cluster scheduling software
- Experience managing HPC low-latency, high-bandwidth interconnects
- Experience supporting Linux-based scientific workstations running visualization applications