Sr. High Performance Computing (HPC) Systems Engineer

SpaceX

Pennsylvania

On-site

USD 150,000 - 190,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

SpaceX is seeking a Senior High Performance Computing (HPC) Systems Engineer to join the HPC team. The role involves administering HPC clusters, storage systems, and high‑speed networks, providing application support, and integrating Linux compute clusters.

The candidate should demonstrate strong Linux, Kubernetes, Python/ Bash scripting, and documentation skills. Preferred experience includes HPC deployment, cluster management with Slurm/PBS/LSF, monitoring with Prometheus/Grafana/Nagios, and

Qualifications

  • Bachelor’s degree in computer science, engineering, math, or scientific discipline and 5+ years of systems engineering experience; OR 7+ years of professional experience building software in lieu of a degree
  • 5+ years of hands-on experience with client and server hardware/software, management tools, enterprise networking, virtualization, and security technologies
  • Experience with Kubernetes
  • Experience with Linux systems, scripting (Bash/Python) and HPC clustering environments
  • Familiarity with cluster resource managers (Slurm, PBS, LSF)
  • Experience with monitoring and alerting technologies (Prometheus, Grafana, Nagios)
  • Familiarity with GPU usage in compute clusters and CUDA
  • Experience deploying automation/configuration management (Puppet, Ansible)
  • ITAR requirements and ability to obtain TS/SCI with Polygraph

Responsibilities

  • Administer and manage HPC clusters, storage systems, and high-speed networks
  • Provide application support to SpaceX personnel across engineering disciplines
  • Install and integrate Linux-based compute clusters
  • Write instructional documentation and convey highly technical ideas in non-technical terms

Skills

Linux
Kubernetes
Python
Bash
Docker
Documentation
GPU computing

Education

Bachelor’s degree in computer science, engineering, math, or scientific discipline

Tools

Docker
Podman
Singularity
Ansible
Puppet
Slurm
PBS
LSF
Prometheus
Grafana
Nagios

Job description

SpaceX was founded under the belief that a future where humanity is out exploring the stars is fundamentally more exciting than one where we are not. Today SpaceX is actively developing the technologies to make this possible, with the ultimate goal ofenabling human life on Mars.

SR. HIGH PERFORMANCE COMPUTING (HPC) SYSTEMS ENGINEER

SpaceX is looking for an HPC Systems Engineer with strong knowledge and experience in a world class engineering organization. This employee will be a member of the HPC team and will support SpaceX personnel and proprietary systems. The ideal candidate will be flexible and flourish in a fast paced and challenging environment. They should be a self-starter, self-motivator and possess ingenuity to excel at this position.

RESPONSIBILITIES:
  • Administer and manage HPC clusters, storage systems, and high-speed networks
  • Provide application support to SpaceX employees across engineering disciplines
  • Install and integrate Linux-based compute clusters
  • Write instructional documentation and convey highly technical ideas in non-technical terms
BASIC QUALIFICATIONS:
  • Bachelor’s degree in computer science, engineering, math, or scientific discipline and 5+ years of systems engineering experience; OR 7+ years of professional experience building software in lieu of a degree
  • 5+ years of hands-on experience with client and server hardware/software, management tools, enterprise networking, virtualization, and security technologies
  • Experience with Kubernetes
PREFERRED SKILLS AND EXPERIENCE:
  • 5+ years of professional experience building, deploying and troubleshooting Linux systems
  • Experience with a scripting language (Bash, Python) to automate and solve reoccurring tasks
  • Experience building, deploying and troubleshooting HPC clusters
  • Familiarity with cluster resource managers (Slurm, PBS, LSF)
  • Experience with monitoring and alerting technologies (Prometheus, Grafana, Nagios)
  • Familiarity with scientific and engineering computing (CFD, FEA)
  • Familiarity with large scale AI training
  • Familiarity with GPU usage in a compute cluster and Cuda
  • Experience with containers (Docker, Podman, Singularity)
  • Experience deploying and maintaining automated configuration management software (Puppet, Ansible)
  • Comfortable working with mission critical and sensitive systems, with a sense of urgency appropriate to the responsibilities
  • Eligibility for access to classified material up to TS/SCI with Polygraph
ADDITIONAL REQUIREMENTS:
  • Must be willing to work extended hours and weekends as needed
ITAR REQUIREMENTS:
  • To conform to U.S. Government export regulations, applicant must be a (i) U.S. citizen or national, (ii) U.S. lawful, permanent resident (aka green card holder), (iii) Refugee under 8 U.S.C. § 1157, or (iv) Asylee under 8 U.S.C. § 1158, or be eligible to obtain the required authorizations from the U.S. Department of State. Learn more about the ITAR here.

SpaceX is an Equal Opportunity Employer; employment with SpaceX is governed on the basis of merit, competence and qualifications and will not be influenced in any manner by race, color, religion, gender, national origin/ethnicity, veteran status, disability status, age, sexual orientation, gender identity, marital status, mental or physical disability or any other legally protected status.

Applicants wishing to view a copy of SpaceX’s Aff… or applicants requiring reasonable accommodation to the application/interview process should reach out to EEOCompliance@spacex.com.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Sr. High Performance Computing (HPC) Systems Engineer
Sr. High Performance Computing (HPC) Systems Engineer

InvestedintheMission • Town of Texas (WI)

On-site
USD 120,000 - 190,000
Sr. High Performance Computing (HPC) Systems Engineer
Sr. High Performance Computing (HPC) Systems Engineer

Future Ventures • Town of Texas (WI)

On-site
USD 140,000 - 210,000
Sr. High Performance Computing (HPC) Systems Engineer
Sr. High Performance Computing (HPC) Systems Engineer

SpaceX • Brownsville (TX)

On-site
USD 140,000 - 200,000
Sr. High Performance Computing (HPC) Systems Engineer
Sr. High Performance Computing (HPC) Systems Engineer

InvestedintheMission • Hawthorne (CA)

On-site
USD 165,000 - 230,000
Stock options
Long-term incentives
Bonuses
+6
Sr. High Performance Computing (HPC) Systems Engineer
Sr. High Performance Computing (HPC) Systems Engineer

SpaceX • Hawthorne (CA)

On-site
USD 165,000 - 230,000
Stock options
Health, vision and dental coverage
401(k) retirement plan
+2
Sr. High Performance Computing (HPC) Systems Engineer
Sr. High Performance Computing (HPC) Systems Engineer

SPACE EXPLORATION TECHNOLOGIES CORP • United States

On-site
USD 140,000 - 210,000
Sr. High Performance Computing (HPC) Systems Engineer
Sr. High Performance Computing (HPC) Systems Engineer

SPACE EXPLORATION TECHNOLOGIES CORP • Hawthorne (CA)

On-site
USD 165,000 - 230,000
Stock options
Medical Vision Dental
Site Reliability Engineer - HPC & Automation (Silicon Engineering)
Site Reliability Engineer - HPC & Automation (Silicon Engineering)

SpaceX • Union Hill-Novelty Hill (WA)

On-site
USD 125,000 - 150,000
Comprehensive medical, vision, and dental coverage
401(k) retirement plan
Paid parental leave
+2
Site Reliability Engineer — HPC & Automation (Silicon Engineering)
Site Reliability Engineer — HPC & Automation (Silicon Engineering)

InvestedintheMission • Redmond (WA)

On-site
USD 125,000 - 175,000
Stock options
401(k) plan
Medical, vision, dental coverage
+2
Site Reliability Engineer — HPC & Automation (Silicon Engineering)
Site Reliability Engineer — HPC & Automation (Silicon Engineering)

SpaceX • Redmond (WA)

On-site
USD 125,000 - 150,000
Comprehensive medical, vision, and dental coverage
401(k) retirement plan
Paid parental leave
+2