High-Performance Computing (HPC) Engineer

aero

El Segundo (CA)

On-site

USD 140,000 - 190,000

Full time

3 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

The Aerospace Corporation is seeking an HPC Engineer/Site Reliability Engineer to lead design and operation of 10,000-core classified and 5,000-core unclassified HPC clusters. You will collaborate with scientists, implement automation with Clush, and apply IaC/GitOps to optimize performance in both classified and unclassified environments.

Requires a Bachelor's in CS/Engineering, 7+ years Linux/HPC experience, Slurm management, CUDA GPU tech, and TS/SCI eligibility.

Qualifications

  • Bachelor's degree in Computer Science, Engineering, or equivalent experience.
  • 7+ years in Linux system administration within an enterprise HPC environment.
  • Experience with environment modules and technical software.
  • Strong knowledge of Linux, networking, and HPC systems.
  • Experience with Infrastructure-as-Code and GitOps.
  • Proven experience managing Slurm and HPC workloads (interactive and batch).
  • Experience with AI & NVIDIA GPUs (CUDA).
  • Scripting and automation with Clush.
  • Experience hardening Linux for security requirements.
  • Experience with hardware/infrastructure automation with vendors like HPE or Cisco.
  • Excellent communication and teamwork across distributed teams.
  • DoD 8570.01-m IAT Level II equivalent; TS/SCI eligible.

Responsibilities

  • Collaborate with scientists and engineers on mission-critical analyses.
  • Lead cross-functional teams and mentor junior engineers.
  • Design and implement HPC solutions for diverse workloads.
  • Manage on-premise and cloud HPC clusters for peak performance.
  • Deliver HPC infrastructure design and configuration.
  • Develop automation using tools such as Clush.
  • Apply GitOps and IaC to manage infrastructure.
  • Implement, support, and optimize GPU computing.
  • Monitor and tune HPC system performance and resource use.
  • Develop cost-efficient HPC service offerings aligned with missions.
  • Harden Linux systems to meet security requirements.

Skills

Linux administration
HPC systems
Slurm scheduler
Infrastructure as Code
GitOps
CUDA
Scripting
Team leadership
Security+
TS/SCI clearance

Education

Bachelor's degree in Computer Science or related field

Tools

Clush
Slurm
Kubernetes
Nsight

Job description

The Aerospace Corporation is the trusted partner to the nation's space programs, solving the hardest problems and providing unmatched technical expertise. As the operator of a federally funded research and development center (FFRDC), we are broadly engaged across all aspects of space- delivering innovative solutions that span satellite, launch, ground, and cyber systems for defense, civil and commercial customers. When you join our team, you'll be part of a special collection of problem solvers, thought leaders, and innovators. Join us and take your place in space.

The Aerospace Corporation is seeking a talented High-Performance Computing (HPC) Engineer ( Site Reliability Engineer Staff III/IV ) to join our Computational Services team. In this role, you will develop, implement, and optimize HPC clusters that support both on-premises and cloud environments. You will work alongside rocket scientists and engineers, tackling complex space enterprise challenges while having direct impact on critical national security missions. We value a collaborative, proactive mindset and a shared commitment to engineering excellence.

The selected candidate will be required to work full-time, on-site at our facility in El Segundo, CA or Chantilly, VA.

What You'll Be Doing
  • Collaborate with scientists and engineers on diverse projects supporting mission-critical technical analysis for national space assets
  • Lead cross-functional teams and mentor junior engineers.
  • Design and implement HPC solutions that optimize resource utilization across diverse workloads in both classified and unclassified settings.
  • Manage on premise 10,000-core classified cluster and a 5,000-core unclassified cluster to ensure peak performance.
  • Deliver high-quality HPC infrastructure design, and system configuration.
  • Develop and deploy automation solutions using tools such as Clush.
  • Manage infrastructure using Infrastructure-as-Code and GitOps practices
  • Implement, support, and optimize GPU computing.
  • Monitor, analyze, and tune HPC system performance, utilization, and resource allocation to maintain operational efficiency.
  • Develop cost-efficient HPC service offerings that align with mission and business objectives.
  • Harden Linux systems to meet stringent security requirements
What You Need to be Successful

Minimum Requirements for the Site Reliability Engineer Staff III :

  • Bachelor's degree in Computer Science, Engineering, or equivalent experience.
  • Minimum of 7 years' experience in Linux system administration within an enterprise HPC environment.
  • Experience supporting technical software (compilers, mod&sim tools, languages, COTS, GOTs) including the development of environment modules.
  • In-depth knowledge of Linux, networking, and HPC systems.
  • Experience with Infrastructure-as-Code and GitOps
  • Proven experience in managing the Slurm scheduler and setting up HPC systems for both interactive and batch workloads.
  • Experience provisioning and supporting AI & NVIDIA GPU technologies (e.g. CUDA)
  • Proficiency in scripting and competence with automation tools such as Clush.
  • Experience hardening Linux systems to meet security requirements
  • Experience with hardware and infrastructure automation in environments using server vendors such as HPE or Cisco.
  • Strong communication skills, with an ability to work both independently and as part of a geographically distributed team.
  • CompTIA Security+ CE certification or equivalent that meets DoD 8570.01-m requirements for IAT Level II personnel
  • Ability to obtain and maintain a TS/SCI clearance (U.S. citizenship required).
  • Demonstrated ability to lead cross-functional teams and mentor junior engineers.

In addition to the above, the minimum requirements for the Site Reliability Engineer Staff IV include:

  • 9+ years of experience in an enterprise 100+ server HPC cluster operations and administration
  • Experience with performance analysis and optimization with custom developed technical software in collaboration with scientists and engineers.
  • Expertise in optimizing and customizing Slurm partitions, qualities of service and priority to balance utilization and reduce job wait times.
  • Experience performing in-place upgrades of Slurm.
  • Implementing visualization of live system telemetry
  • Experience developing and architecting cluster configuration management
  • Advanced Infrastructure-as-Code GitOps (e.g. multi-branch pipelines)
  • Skill in provisioning and supporting AI & NVIDIA GPU technologies (e.g. CUDA, Nsight), with expertise in GPU integration, resource allocation, and scheduling using Slurm.
How You Can Stand Out

It would be impressive if you have one or more of these:

  • An active TS/SCI clearance with CI Polygraph.
  • Experience integrating Slurm with SELinux
  • Experience implementing DISA STIG compliance
  • Experience integrating Kubernetes and Slurm, e.g. Slinky
  • Experience supporting and managing diverse HPC workloads, including computational fluid dynamics, Monte Carlo, structural analysis
  • Experience integrating user web portals to launch & manage wor
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

High-Performance Computing (HPC) Engineer
High-Performance Computing (HPC) Engineer

The Aerospace Corporation • El Segundo (CA)

On-site
USD 135,000 - 203,000
Health care benefits
Paid time off
Telework options
Space HPC Engineer & SRE - On-Site, TS/SCI Ready
Space HPC Engineer & SRE - On-Site, TS/SCI Ready

The Aerospace Corporation • El Segundo (CA)

On-site
USD 135,000 - 203,000
Health care benefits
Paid time off
Telework options
Space HPC Engineer & SRE (Slurm, GPU, Cloud)
Space HPC Engineer & SRE (Slurm, GPU, Cloud)

aero • El Segundo (CA)

On-site
USD 140,000 - 190,000
HPC Infrastructure and Cluster Engineer
HPC Infrastructure and Cluster Engineer

Arena Technical Resources, LLC (ATR) • Springfield (VA)

On-site
USD 180,000 - 200,000
Classified Windows System Administrator
Classified Windows System Administrator

The Aerospace Corporation • Chantilly (VA)

On-site
USD 75,000 - 113,000
HPC Linux Systems Engineer
HPC Linux Systems Engineer

Cadre5 • Knoxville (TN)

Hybrid
USD 120,000 - 160,000
Excellent medical insurance
Employer-paid benefits
HPC Software Engineer
HPC Software Engineer

Cornerstone Defense LLC • Colorado Springs (CO)

On-site
USD 140,000 - 230,000
Sr. High Performance Computing (HPC) Systems Engineer
Sr. High Performance Computing (HPC) Systems Engineer

SpaceX • Town of Texas (WI)

On-site
USD 140,000 - 190,000
Linux Systems Administrator – Top Secret HPC
Linux Systems Administrator – Top Secret HPC

GIGATEC Engineering • Maryland

On-site
Senior HPC DevOps Engineer | TS/SCI w/ MD POLY Security Clearance required
Senior HPC DevOps Engineer | TS/SCI w/ MD POLY Security Clearance required

Capstone Technology Partners • College Park (MD)

On-site
USD 222,000 - 257,000
Four weeks paid time off
Eleven paid holidays
401k with employer contributions and 3
+2