Sr. High Performance Computing (HPC) Systems Engineer

United States Digital Space LLC

Starbase (TX)

On-site

USD 120,000 - 190,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

United States Digital Space LLC seeks an Sr. HPC Systems Engineer to administrate HPC clusters, storage, and fast networks. You will support engineers, install Linux compute clusters, and author clear documentation for complex systems.

Ideal candidates have 5+ years in HPC or 7+ years software experience with a degree, plus Kubernetes and container tooling proficiency. This role involves extended hours as needed.

Qualifications

  • Bachelor's degree in computer science, engineering, math, or scientific discipline and 5+ years of systems engineering experience; OR 7+ years of professional experience building software in lieu of a degree
  • 5+ years of hands-on experience with client and server hardware/software, management tools, enterprise networking, virtualization, and security technologies
  • Experience with Kubernetes

Responsibilities

  • Administer and manage HPC clusters, storage systems, and high-speed networks
  • Provide application support to the company employees across engineering disciplines
  • Install and integrate Linux-based compute clusters
  • Write instructional documentation and convey highly technical ideas in non-technical terms

Skills

Kubernetes
Linux systems
Scripting Bash Python
Docker Podman Singularity
HPC clusters
Monitoring Prometheus Grafana Nagios
Cluster schedulers Slurm PBS LSF
Automation Puppet Ansible
Virtualization
Networking security
GPU CUDA

Education

Bachelor's degree in computer science, engineering, math, or scientific discipline

Tools

Docker
Podman
Singularity
Kubernetes
Puppet
Ansible
Prometheus
Grafana
Nagios
Slurm
PBS
LSF

Job description

the company was founded under the belief that a future where humanity is out exploring the stars is fundamentally more exciting than one where we are not. Today the company is actively developing the technologies to make this possible, with the ultimate goal of enabling human life on Mars.

SR. HIGH PERFORMANCE COMPUTING (HPC) SYSTEMS ENGINEER

the company is looking for an HPC Systems Engineer with strong knowledge and experience in a world class engineering organization. This employee will be a member of the HPC team and will support the company personnel and proprietary systems. The ideal candidate will be flexible and flourish in a fast paced and challenging environment. They should be a self-starter, self-motivator and possess ingenuity to excel at this position.

RESPONSIBILITIES
  • Administer and manage HPC clusters, storage systems, and high-speed networks
  • Provide application support to the company employees across engineering disciplines
  • Install and integrate Linux-based compute clusters
  • Write instructional documentation and convey highly technical ideas in non-technical terms
BASIC QUALIFICATIONS
  • Bachelor's degree in computer science, engineering, math, or scientific discipline and 5+ years of systems engineering experience; OR 7+ years of professional experience building software in lieu of a degree
  • 5+ years of hands-on experience with client and server hardware/software, management tools, enterprise networking, virtualization, and security technologies
  • Experience with Kubernetes
PREFERRED SKILLS AND EXPERIENCE
  • 5+ years of professional experience building, deploying and troubleshooting Linux systems
  • Experience with a scripting language (Bash, Python) to automate and solve reoccurring tasks
  • Experience building, deploying and troubleshooting HPC clusters
  • Familiarity with cluster resource managers (Slurm, PBS, LSF)
  • Experience with monitoring and alerting technologies (Prometheus, Grafana, Nagios)
  • Familiarity with scientific and engineering computing (CFD, FEA)
  • Familiarity with large scale AI training
  • Familiarity with GPU usage in a compute cluster and Cuda
  • Experience with containers (Docker, Podman, Singularity)
  • Experience deploying and maintaining automated configuration management software (Puppet, Ansible)
  • Comfortable working with mission critical and sensitive systems, with a sense of urgency appropriate to the responsibilities
  • Eligibility for access to classified material up to TS/SCI with Polygraph
ADDITIONAL REQUIREMENTS
  • Must be willing to work extended hours and weekends as needed
ITAR REQUIREMENTS
  • To conform to U.S. Government export regulations, applicant must be a (i) U.S. citizen or national, (ii) U.S. lawful, permanent resident (aka green card holder), (iii) Refugee under 8 U.S.C. Section 1157, or (iv) Asylee under 8 U.S.C. Section 1158, or be eligible to obtain the required authorizations from the U.S. Department of State. Learn more about the ITAR here.

the company is an Equal Opportunity Employer; employment with the company is governed on the basis of merit, competence and qualifications and will not be influenced in any manner by race, color, religion, gender, national origin/ethnicity, veteran status, disability status, age, sexual orientation, gender identity, marital status, mental or physical disability or any other legally protected status.

Applicants wishing to view a copy of the company’s Aff... or applicants requiring reasonable accommodation to the application/interview process should reach out to hr@unitedstatesdigital.space.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Sr. High Performance Computing (HPC) Systems Engineer
Sr. High Performance Computing (HPC) Systems Engineer

United States Digital Space LLC • Hawthorne (CA)

On-site
USD 165,000 - 230,000
Sr. High Performance Computing (HPC) Systems Engineer
Sr. High Performance Computing (HPC) Systems Engineer

InvestedintheMission • Town of Texas (WI)

On-site
USD 120,000 - 190,000
Sr. High Performance Computing (HPC) Systems Engineer
Sr. High Performance Computing (HPC) Systems Engineer

SpaceX • Pennsylvania

On-site
USD 150,000 - 190,000
Sr. High Performance Computing (HPC) Systems Engineer
Sr. High Performance Computing (HPC) Systems Engineer

SpaceX • Brownsville (TX)

On-site
USD 140,000 - 200,000
Sr. High Performance Computing (HPC) Systems Engineer
Sr. High Performance Computing (HPC) Systems Engineer

Future Ventures • Town of Texas (WI)

On-site
USD 140,000 - 210,000
Sr. High Performance Computing (HPC) Systems Engineer
Sr. High Performance Computing (HPC) Systems Engineer

SpaceX • Hawthorne (CA)

On-site
USD 165,000 - 230,000
Stock options
Health, vision and dental coverage
401(k) retirement plan
+2
Sr. High Performance Computing (HPC) Systems Engineer
Sr. High Performance Computing (HPC) Systems Engineer

SPACE EXPLORATION TECHNOLOGIES CORP • United States

On-site
USD 140,000 - 210,000
Sr. High Performance Computing (HPC) Systems Engineer
Sr. High Performance Computing (HPC) Systems Engineer

InvestedintheMission • Hawthorne (CA)

On-site
USD 165,000 - 230,000
Stock options
Long-term incentives
Bonuses
+6
Sr. IT Systems Administrator (Top Secret Clearance)
Sr. IT Systems Administrator (Top Secret Clearance)

United States Digital Space LLC • Washington

On-site
USD 95,000 - 150,000
Comprehensive medical, vision, and dental coverage
401(k) retirement plan
Paid parental leave
+2
Site Reliability Engineer — HPC & Automation (Silicon Engineering)
Site Reliability Engineer — HPC & Automation (Silicon Engineering)

United States Digital Space LLC • Redmond (WA)

On-site
USD 125,000 - 175,000
Comprehensive medical, vision, and dental coverage
401(k) retirement plan
Paid vacation and holidays