Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.
The Aerospace Corporation in El Segundo, CA is seeking an HPC Engineer (Site Reliability Engineer Staff III/IV) to join our Computational Services team. You will design, deploy, and optimize HPC clusters on prem and in cloud, collaborating with rocket scientists and engineers on mission-critical analyses.
Minimum qualifications include a BS in CS/engineering, 7+ years Linux admin in HPC, Slurm expertise, and experience with IaC, GitOps, and AI/NVIDIA GPUs.
The Aerospace Corporation is the trusted partner to the nation’s space programs, solving the hardest problems and providing unmatched technical expertise. As the operator of a federally funded research and development center (FFRDC), we are broadly engaged across all aspects of space— delivering innovative solutions that span satellite, launch, ground, and cyber systems for defense, civil and commercial customers. When you join our team, you’ll be part of a special collection of problem solvers, thought leaders, and innovators. Join us and take your place in space. The Aerospace Corporation is seeking a talented High-Performance Computing (HPC) Engineer (Site Reliability Engineer Staff III/IV) to join our Computational Services team. In this role, you will develop, implement, and optimize HPC clusters that support both on-premises and cloud environments. You will work alongside rocket scientists and engineers, tackling complex space enterprise challenges while having direct impact on critical national security missions. We value a collaborative, proactive mindset and a shared commitment to engineering excellence. The selected candidate will be required to work full-time, on-site at our facility in El Segundo, CA or Chantilly, VA.
Collaborate with scientists and engineers on diverse projects supporting mission‑critical technical analysis for national space assets Lead cross‑functional teams and mentor junior engineers. Design and implement HPC solutions that optimize resource utilization across diverse workloads in both classified and unclassified settings. Manage on premise 10,000-core classified cluster and a 5,000-core unclassified cluster to ensure peak performance. Deliver high‑quality HPC infrastructure design, and system configuration. Develop and deploy automation solutions using tools such as Clush. Manage infrastructure using Infrastructure‑as‑Code and GitOps practices Implement, support, and optimize GPU computing. Monitor, analyze, and tune HPC system performance, utilization, and resource allocation to maintain operational efficiency. Develop cost‑efficient HPC service offerings that align with mission and business objectives. Harden Linux systems to meet stringent security requirements.
Minimum Requirements for the Site Reliability Engineer Staff III: Bachelor’s degree in Computer Science, Engineering, or equivalent experience. Minimum of 7 years’ experience in Linux system administration within an enterprise HPC environment. Experience supporting technical software (compilers, mod&sim tools, languages, COTS, GOTs) including the development of environment modules. In‑depth knowledge of Linux, networking, and HPC systems. Experience with Infrastructure‑as‑Code and GitOps Proven experience in managing the Slurm scheduler and setting up HPC systems for both interactive and batch workloads. Experience provisioning and supporting AI & NVIDIA GPU technologies(e.g. CUDA) Proficiency in scripting and competence with automation tools such as Clush. Experience hardening Linux systems to meet security requirements Experience with hardware and infrastructure automation in environments using server vendors such as HPE or Cisco. Strong communication skills, with an ability to work both independently and as part of a geographically distributed team. CompTIA Security+ CE certification or equivalent that meets DoD 8570.01‑m requirements for IAT Level II personnel Ability to obtain and maintain a TS/SCI clearance (U.S. citizenship required). Demonstrated ability to lead cross‑functional teams and mentor junior engineers.
In addition to the above, the minimum requirements for the Site Reliability Engineer Staff IV include: 9+ years of experience in an enterprise 100+ server HPC cluster operations and administration Experience with performance analysis and optimization with custom developed technical software in collaboration with scientists and engineers. Expertise in optimizing and customizing Slurm partitions, qualities of service and priority to balance utilization and reduce job wait times. Experience performing in‑place upgrades of Slurm. Implementing visualization of live system telemetry Experience developing and architecting cluster configuration management Advanced Infrastructure‑as‑Code GitOps (e.g. multi‑branch pipelines) Skill in provisioning and supporting AI & NVIDIA GPU technologies (e.g. CUDA, Nsight), with expertise in GPU integration, resource allocation, and scheduling using Slurm.
It would be impressive if you have one or more of these: An active TS/SCI clearance with CI Polygraph. Experience integrating Slurm with SELinux Experience implementing DISA STIG compliance Experience integrating Kubernetes and Slurm, e.g. Slinky Experience supporting and managing diverse HPC workloads, including computational fluid dynamics, Monte Carlo, structural analysis Experience integrating user web portals to launch & manage workloads (e.g. Open OnDemand and developing plugins for session persistence and VSCode) Experience implementing utilization dashboards (e.g. XDMod) with Slurm Experience managing parallel file systems such as Lustre. Experience developing solutions that optimize data storage Experience implementing or supporting Slurm REST API Experience with automated provisioning, e.g. Warewolf, Kickstart, PXE Knowledge of NVLINK and DCGM for optimizing GPU workflows. Familiarity with Prometheus and Grafana for monitoring and performance visualization. Background in containerization within an HPC context usedfor data processing and technical analysis Experience packaging custom software (e.g. RPMs) Proficiency with automation tools such as Ansible for HPC Hands‑on background with cloud HPC services Experience with AWS Parallel Computing Service (AWS ParallelCluster).
Leadership Competencies Our leadership philosophy is simple: every employee, regardless of level and role, can demonstrate leadership. At Aerospace, our commitment is our people. To cultivate our talent and ensure that we have a strong pipeline of future leaders, we want individuals who: Operate Strategically Lead Change Engage with Impact Foster Innovation Deliver Results Ways We Reward Our Employees During your interview process, our team will provide details of our industry‑leading benefits.
We are all unique, from various backgrounds and all walks of life, yet one thing bonds all of us to each other—the belief that we can make a difference. This core belief empowers us to do our best work at The Aerospace Corporation.
The Aerospace Corporation is an equal opportunity employer. All qualified applicants will receive consideration for employment and will not be discriminated against on the basis of race, age, sex (including pregnancy, childbirth, and related medical conditions), sexual orientation, gender, gender identity or expression, color, religion, genetic information, marital status, ancestry, national origin, protected veteran status, physical disability, medical condition, mental disability, or disability status and any other characteristic protected by state or federal law. If you’re an individual with a disability or a disabled veteran who needs assistance using our online job search and application tools or need reasonable accommodation to complete the job application process, please contact us by phone at 310.336.5432 or by email at peoplemangmnt.mailbox@aero.org . You can also review Know Your Rights: Workplace Discrimination is Illegal.
The Aerospace Corporation has over 26 locations nationwide, which allows us to build a workforce that will nurture the best ideas, and generate the most novel innovations and solutions to our nation’s toughest space enterprise challenges.
Aerospace covers all stages of the space lifecycle, from concept to operations. We believe our people are our most valuable resource and that’s why we propel our employees forward in their careers through ongoing education, training, and professional development.
We can all serve a mission much greater than ourselves, and at Aerospace that is exactly what we do. By hiring the industry’s most preeminent scientists and engineers, Aerospace advances emerging technologies that protect some of the most critical missions on Earth and above it.
Aerospace employees working in organizations with technical responsibilities are required to obtain a Security Clearance. U.S. citizenship is required for those positions.