and promotes a high level of staff retention. We offer services ranging from full life cycle HPC systems engineering to remote managed services to HPC program analysis.
RedLine is looking for a Senior High Performance Computing (HPC) Administrator to join our team. This position will work on various on-premise installations in support of new flagship supercomputer and managed infrastructure located at Lawrence Berkeley National Laboratory. The ideal candidate will be an experienced individual with a strong security, Linux, HPC, configuration management, systems automation and networking background. This position supports requires strong organization, management and customer facing skills.
The customer supports a high-visibility, next-generation national computing initiative supporting cutting-edge scientific research and discovery. As a Sr. HPC System Administrator, you will directly contribute to the reliability, performance, and availability of the advanced computing environment that enables large-scale scientific research and breakthroughs with far-reaching impact.
US citizenship and the ability to obtain a Public Trust clearance is a requirement to apply. This is a remote position but will work Pacific Time core hours. This full-time position offers a full benefits package including paid time off, 401k match, and health care benefits.
Job Responsibilities:
Lead a team to administer resources from the operating system and above within on-premise HPC environments. Efforts include, but are not limited to:
- Integration and configuration of compute resources and service nodes
- All software installations on compute resources, services nodes, and parallel filesystems
- Develop/implement system and performance monitoring and benchmarking
- Maintain system documentation
- Evaluate performance impacts of planned operating system changes
- Lead resource optimization and job scheduling software and policies
- Provide technical support to researchers using HPC resources, troubleshoot problems and develop appropriate computational strategies
- Provide emergency support on a 24x7 basis.
- Provide technical leadership and direction for other team members. Maintain team focus on production uptime and model performance.
- Review and present all change management requests to customer management.
- Work both independently and as part of the team; able to concurrently work on several projects.
- Effectively communicate with people of diverse backgrounds and computer knowledge.
- Manage individual and team task pipelines, ensuring all deliverables are met and tracking systems are consistently updated.
- Solicit and analyze customer feedback, ensuring critical issues are effectively communicated to the team and captured as trackable deliverables.
- Manage individual and team task pipelines, ensuring all deliverables are met and tracking systems are consistently updated.
- Solicit and analyze customer feedback, ensuring critical issues are effectively communicated to the team and captured as trackable deliverables.
- Travel to the client site in California approximately 1 week per quarter with some additional travel upon installation of the system.
Requirements:
- Minimum of 10 years RedHat, Rocky, and/or CentOS Linux system administrator experience.
- Technical leadership experience in a large, production environment – leadership for the technical solution and the system administration team.
- Demonstrated ability to configure, deploy and manage major system areas such as batch system, network, data storage, backup system, database system, or distributed computing
- Experience with configuration management tools (e.g., Ansible)
- Ability to work both independently and as part of the team; flexibility in dealing with assignments and in working on several projects simultaneously
- Ability to effectively communicate with people of diverse backgrounds and computer knowledge.
Preferred Skills:
- HPC system administration experience is highly preferred.
- Experience with batch systems (e.g., SLURM)
- Experience managing parallel and cluster file systems (e.g., Lustre)
- Network management experience, including in an HPC context (e.g., InfiniBand)