Space HPC Engineer & SRE (Slurm, GPU, Cloud)

aero

El Segundo (CA)

On-site

USD 140,000 - 190,000

Full time

3 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

The Aerospace Corporation is seeking an HPC Engineer/Site Reliability Engineer to lead design and operation of 10,000-core classified and 5,000-core unclassified HPC clusters. You will collaborate with scientists, implement automation with Clush, and apply IaC/GitOps to optimize performance in both classified and unclassified environments.

Requires a Bachelor's in CS/Engineering, 7+ years Linux/HPC experience, Slurm management, CUDA GPU tech, and TS/SCI eligibility.

Qualifications

  • Bachelor's degree in Computer Science, Engineering, or equivalent experience.
  • 7+ years in Linux system administration within an enterprise HPC environment.
  • Experience with environment modules and technical software.
  • Strong knowledge of Linux, networking, and HPC systems.
  • Experience with Infrastructure-as-Code and GitOps.
  • Proven experience managing Slurm and HPC workloads (interactive and batch).
  • Experience with AI & NVIDIA GPUs (CUDA).
  • Scripting and automation with Clush.
  • Experience hardening Linux for security requirements.
  • Experience with hardware/infrastructure automation with vendors like HPE or Cisco.
  • Excellent communication and teamwork across distributed teams.
  • DoD 8570.01-m IAT Level II equivalent; TS/SCI eligible.

Responsibilities

  • Collaborate with scientists and engineers on mission-critical analyses.
  • Lead cross-functional teams and mentor junior engineers.
  • Design and implement HPC solutions for diverse workloads.
  • Manage on-premise and cloud HPC clusters for peak performance.
  • Deliver HPC infrastructure design and configuration.
  • Develop automation using tools such as Clush.
  • Apply GitOps and IaC to manage infrastructure.
  • Implement, support, and optimize GPU computing.
  • Monitor and tune HPC system performance and resource use.
  • Develop cost-efficient HPC service offerings aligned with missions.
  • Harden Linux systems to meet security requirements.

Skills

Linux administration
HPC systems
Slurm scheduler
Infrastructure as Code
GitOps
CUDA
Scripting
Team leadership
Security+
TS/SCI clearance

Education

Bachelor's degree in Computer Science or related field

Tools

Clush
Slurm
Kubernetes
Nsight

Job description

The Aerospace Corporation is seeking an HPC Engineer/Site Reliability Engineer to lead design and operation of 10,000-core classified and 5,000-core unclassified HPC clusters. You will collaborate with scientists, implement automation with Clush, and apply IaC/GitOps to optimize performance in both classified and unclassified environments.

Requires a Bachelor's in CS/Engineering, 7+ years Linux/HPC experience, Slurm management, CUDA GPU tech, and TS/SCI eligibility.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Space HPC Engineer & SRE - On-Site, TS/SCI Ready
Space HPC Engineer & SRE - On-Site, TS/SCI Ready

The Aerospace Corporation • El Segundo (CA)

On-site
USD 135,000 - 203,000
Health care benefits
Paid time off
Telework options
High-Performance Computing (HPC) Engineer
High-Performance Computing (HPC) Engineer

aero • El Segundo (CA)

On-site
USD 140,000 - 190,000
Senior HPC Systems Engineer for High-Performance Clusters
Senior HPC Systems Engineer for High-Performance Clusters

SPACE EXPLORATION TECHNOLOGIES CORP • United States

On-site
USD 140,000 - 210,000
Senior HPC Systems Engineer – Clusters & AI Compute
Senior HPC Systems Engineer – Clusters & AI Compute

SPACE EXPLORATION TECHNOLOGIES CORP • Hawthorne (CA)

On-site
USD 165,000 - 230,000
Stock options
Medical Vision Dental
HPC Infrastructure and Cluster Engineer
HPC Infrastructure and Cluster Engineer

Arena Technical Resources, LLC (ATR) • Springfield (VA)

On-site
USD 180,000 - 200,000
Senior HPC Engineer, Classified Computing Lead
Senior HPC Engineer, Classified Computing Lead

Oak Ridge National Laboratory • Oak Ridge (TN)

On-site
USD 120,000 - 160,000
High-Performance Computing (HPC) Engineer
High-Performance Computing (HPC) Engineer

The Aerospace Corporation • El Segundo (CA)

On-site
USD 135,000 - 203,000
Health care benefits
Paid time off
Telework options
Senior HPC Systems Engineer: Scale AI Clusters
Senior HPC Systems Engineer: Scale AI Clusters

SpaceX • Hawthorne (CA)

On-site
USD 165,000 - 230,000
Stock options
Health, vision and dental coverage
401(k) retirement plan
+2
Senior HPC & GPU Systems Engineer
Senior HPC & GPU Systems Engineer

Xcelerate-Solutions-5 • Bethesda (MD)

On-site
USD 120,000 - 160,000
HPC Cluster Engineer for AI Workloads | TS/SCI
HPC Cluster Engineer for AI Workloads | TS/SCI

Socket.dev • Springfield (VA)

On-site
USD 148,000 - 179,000
Health/Dental/Vision
401(k)
Paid Time Off
+2