High-Performance Computing (HPC) Engineer

The Aerospace Corporation

United States

On-site

USD 135,000 - 203,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Comprehensive health care
401(k) plan
Relocation assistance
Education assistance

Job summary

The Aerospace Corporation seeks a High-Performance Computing (HPC) Engineer (Site Reliability Engineer Staff III/IV) to join our Computational Services team. You will design, implement, and optimize on-premises and cloud HPC clusters while collaborating with scientists and engineers on mission-critical analysis.

The role requires advanced Linux administration, Slurm knowledge, and automation experience; TS/SCI clearance is preferred.

Qualifications

  • Bachelor’s degree in Computer Science, Engineering, or equivalent experience.
  • 7+ years Linux system administration in an enterprise HPC environment.
  • Experience with Slurm scheduler and environment modules.
  • Experience with Infrastructure-as-Code and GitOps practices.
  • Experience with AI & GPU technologies (CUDA) and scripting.

Responsibilities

  • Collaborate with scientists and engineers on mission-critical technical analysis for national space assets.
  • Lead cross-functional teams and mentor junior engineers.
  • Design and optimize HPC clusters for on-premises and cloud environments.
  • Manage 10,000-core classified and 5,000-core unclassified clusters.
  • Develop automation using tools like Clush and implement IaC and GitOps.
  • Harden Linux systems to meet security requirements.

Skills

Linux administration
Slurm scheduler
Automation scripting
GitOps
Clush
HPC design
Communication skills
Security+ / DoD 8570

Education

Bachelor’s degree in Computer Science, Engineering, or equivalent experience

Tools

Environment modules
CUDA (NVIDIA)
Ansible
Terraform
Open OnDemand

Job description

The Aerospace Corporation is the trusted partner to the nation’s space programs, solving the hardest problems and providing unmatched technical expertise. As the operator of a federally funded research and development center (FFRDC), we are broadly engaged across all aspects of space— delivering innovative solutions that span satellite, launch, ground, and cyber systems for defense, civil and commercial customers. When you join our team, you’ll be part of a special collection of problem solvers, thought leaders, and innovators. Join us and take your place in space.

The Aerospace Corporation is seeking a talented High-Performance Computing (HPC) Engineer (Site Reliability Engineer Staff III/IV) to join our Computational Services team. In this role, you will develop, implement, and optimize HPC clusters that support both on-premises and cloud environments. You will work alongside rocket scientists and engineers, tackling complex space enterprise challenges while having direct impact on critical national security missions. We value a collaborative, proactive mindset and a shared commitment to engineering excellence.

The selected candidate will be required to work full-time, on-site at our facility in El Segundo, CA or Chantilly, VA.

What You’ll Be Doing
  • Collaborate with scientists and engineers on diverse projects supporting mission-critical technical analysis for national space assets
  • Lead cross-functional teams and mentor junior engineers.
  • Design and implement HPC solutions that optimize resource utilization across diverse workloads in both classified and unclassified settings.
  • Manage on premise 10,000-core classified cluster and a 5,000-core unclassified cluster to ensure peak performance.
  • Deliver high-quality HPC infrastructure design, and system configuration.
  • Develop and deploy automation solutions using tools such as Clush.
  • Manage infrastructure using Infrastructure-as-Code and GitOps practices
  • Implement, support, and optimize GPU computing.
  • Monitor, analyze, and tune HPC system performance, utilization, and resource allocation to maintain operational efficiency.
  • Develop cost-efficient HPC service offerings that align with mission and business objectives.
  • Harden Linux systems to meet stringent security requirements
What You Need to be Successful

Minimum Requirements for the Site Reliability Engineer Staff III:

  • Bachelor’s degree in Computer Science, Engineering, or equivalent experience.
  • Minimum of 7 years’ experience in Linux system administration within an enterprise HPC environment.
  • Experience supporting technical software (compilers, mod&sim tools, languages, COTS, GOTs) including the development of environment modules.
  • In-depth knowledge of Linux, networking, and HPC systems.
  • Experience with Infrastructure-as-Code and GitOps
  • Proven experience in managing the Slurm scheduler and setting up HPC systems for both interactive and batch workloads.
  • Experience provisioning and supporting AI & NVIDIA GPU technologies(e.g. CUDA)
  • Proficiency in scripting and competence with automation tools such as Clush.
  • Experience hardening Linux systems to meet security requirements
  • Experience with hardware and infrastructure automation in environments using server vendors such as HPE or Cisco.
  • Strong communication skills, with an ability to work both independently and as part of a geographically distributed team.
  • CompTIA Security+ CE certification or equivalent that meets DoD 8570.01-m requirements for IAT Level II personnel
  • Ability to obtain and maintain a TS/SCI clearance (U.S. citizenship required).
  • Demonstrated ability to lead cross-functional teams and mentor junior engineers.

In addition to the above, the minimum requirements for the Site Reliability Engineer Staff IV include:

  • 9+ years of experience in an enterprise 100+ server HPC cluster operations and administration
  • Experience with performance analysis and optimization with custom developed technical software in collaboration with scientists and engineers.
  • Expertise in optimizing and customizing Slurm partitions, qualities of service and priority to balance utilization and reduce job wait times.
  • Experience performing in-place upgrades of Slurm.
  • Implementing visualization of live system telemetry
  • Experience developing and architecting cluster configuration management
  • Advanced Infrastructure-as-Code GitOps (e.g. multi-branch pipelines)
  • Skill in provisioning and supporting AI & NVIDIA GPU technologies (e.g. CUDA, Nsight), with expertise in GPU integration, resource allocation, and scheduling using Slurm.
How You Can Stand Out

It would be impressive if you have one or more of these:

  • An active TS/SCI clearance with CI Polygraph.
  • Experience integrating Slurm with SELinux
  • Experience implementing DISA STIG compliance
  • Experience integrating Kubernetes and Slurm, e.g. Slinky
  • Experience supporting and managing diverse HPC workloads, including computational fluid dynamics, Monte Carlo, structural analysis
  • Experience integrating user web portals to launch & manage workloads (e.g. Open OnDemand and developing plugins for session persistence and VSCode)
  • Experience implementing utilization dashboards (e.g. XDMod) with Slurm
  • Experience managing parallel file systems such as Lustre.
  • Experience developing solutions that optimize data storage
  • Experience implementing or supporting Slurm REST API
  • Experience with automated provisioning, e.g. Warewolf, Kickstart, PXE
  • Knowledge of NVLINK and DCGM for optimizing GPU workflows.
  • Familiarity with Prometheus and Grafana for monitoring and performance visualization.
  • Background in containerization within an HPC context usedfor data processing and technical analysis
  • Experience packaging custom software (e.g. RPMs)
  • Proficiency with automation tools such as Ansible for HPC
  • Hands-on background with cloud HPC services
  • Experience with AWS Parallel Computing Service (AWS ParallelCluster).

We offer a competitive compensation package where you’ll be rewarded based on your performance and recognized for the value you bring to our business. The grade-based pay range for this job is listed below. Individual salaries within that range are determined through a wide variety of factors including but not limited to education, experience, knowledge and skills.

(Min - Max)

$135,200.00 - $202,800.00 Pay Basis: Annual

Leadership Competencies

Our leadership philosophy is simple: every employee, regardless of level and role, can demonstrate leadership. At Aerospace, our commitment is our people. To cultivate our talent and ensure that we have a strong pipeline of future leaders, we want individuals who:

  • Operate Strategically
  • Lead Change
  • Engage with Impact
  • Foster Innovation
  • Deliver Results
Ways We Reward Our Employees

During your interview process, our team will provide details of our industry-leading benefits.

Benefits vary and are applicable based on Job Type. A few highlights include:

  • Comprehensive health care and wellness plans
  • Paid holidays, sick time, and vacation
  • Standard and alternate work schedules, including telework options
  • 401(k) Plan — Employees receive a total company-paid benefit of 8%, 10%, or 12% of eligible compensation based on years of service and matching contributions; employees are immediately eligible and vested in the plan upon hire
  • Flexible spending accounts
  • Variable pay program for exceptional contributions
  • Relocation assistance
  • Professional growth and development programs to help advance your career
  • Education assistance programs
  • An inclusive work environment built on teamwork, flexibility, and respect

We are all unique, from various backgrounds and all walks of life, yet one thing bonds all of us to each other—the belief that we can make a difference. This core belief empowers us to do our best work at The Aerospace Corporation.

Equal Opportunity Commitment

The Aerospace Corporation is an equal opportunity employer. All qualified applicants will receive consideration for employment and will not be discriminated against on the basis of race, age, sex (including pregnancy, childbirth, and related medical conditions), sexual orientation, gender, gender identity or expression, color, religion, genetic information, marital status, ancestry, national origin, protected veteran status, physical disability, medical condition, mental disability, or disability status and any other characteristic protected by state or federal law.

If you’re an individual with a disability or a disabled veteran who needs assistance using our online job search and application tools or need reasonable accommodation to complete the job application process, please contact us by phone at 310.336.5432 or by email at peoplemangmnt.mailbox@aero.org. You can also review Know Your Rights: Workplace Discrimination is Illegal.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

High-Performance Computing (HPC) Engineer
High-Performance Computing (HPC) Engineer

The Aerospace Corporation • Chantilly (VA)

On-site
USD 135,000 - 203,000
Health care plans
401(k) matching
Relocation assistance
+4
High-Performance Computing (HPC) Engineer
High-Performance Computing (HPC) Engineer

The Aerospace Corporation • El Segundo (CA)

On-site
USD 135,000 - 203,000
Health care benefits
Paid time off
Telework options
Kubernetes Site Reliability Engineer
Kubernetes Site Reliability Engineer

The Aerospace Corporation • United States

On-site
USD 129,000 - 194,000
Comprehensive health care and wellness
401(k) Plan
Relocation assistance
+1
Kubernetes Site Reliability Engineer
Kubernetes Site Reliability Engineer

The Aerospace Corporation • Chantilly (VA)

On-site
USD 129,000 - 194,000
Health care and wellness plans
401(k) with company match
Relocation assistance
+2
Artificial Intelligence and High-Performance Computing Practitioner
Artificial Intelligence and High-Performance Computing Practitioner

aero • Chantilly (VA)

On-site
USD 129,000 - 194,000
Health care
401(k) plan
Relocation assistance
+2
Kubernetes Site Reliability Engineer
Kubernetes Site Reliability Engineer

aero • El Segundo (CA)

On-site
USD 129,000 - 193,500
Comprehensive health care and wellness plans
401(k) Plan with company-paid benefits
Professional growth and development programs
2027 Data Scientist/Machine Learning Engineer
2027 Data Scientist/Machine Learning Engineer

The Aerospace Corporation • El Segundo (CA)

On-site
USD 70,000 - 105,000
Health care and wellness plans
Paid holidays, sick time, and vacation
401(k) plan with employer matching and
+3
Software Integration Engineer
Software Integration Engineer

The Aerospace Corporation • Chantilly (VA)

On-site
USD 86,600 - 129,800
Comprehensive health care and wellness plans
Paid holidays, sick time, and vacation
401(k) Plan with company-paid benefits
Artificial Intelligence and High-Performance Computing Practitioner
Artificial Intelligence and High-Performance Computing Practitioner

The Aerospace Corporation • Chantilly (VA)

On-site
USD 129,000 - 194,000
Health care
Paid holidays
Relocation assistance
+2
Project Management Staff III – International Programs & Enterprise Architect Engineering
Project Management Staff III – International Programs & Enterprise Architect Engineering

The Aerospace Corporation • United States

On-site
USD 85,000 - 127,000
Comprehensive health care and wellness
Paid holidays, sick time, and vacation
Telework options
+1