Senior HPC Systems Engineer: Slurm, GPU & Hybrid Clusters

Parallel Works

United States

Hybrid

USD 180,000 - 230,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Medical, vision, and dental coverage
401(k) with company match
Paid vacation & sick time
Short term disability

Job summary

Parallel Works is seeking a Senior HPC Systems Engineer to design, operate and secure GPU-accelerated clusters behind defense and research programs. You will manage GPU node bring-up, Slurm configuration, fabric and storage troubleshooting, and security hardening across on-prem and government cloud environments.

The role is senior, handling escalations, coordinating with site staff, and mentoring junior engineers while ensuring high availability and compliance in risk-conscious environments.

Qualifications

  • 10+ years operating production Linux systems across multiple distributions (RHEL/ Rocky/ Alma on Gov side; Debian/Ubuntu on commercial GPU side).
  • Experience with Slurm, HPC scheduling, and cluster hardening in a hybrid environment.
  • Proficiency with Bash and Python scripting.
  • DevOps tooling experience with infrastructure as code (Ansible, Terraform).
  • Security and incident response involving STIGs, FIPS cryptography, and secure network configurations.

Responsibilities

  • Operate and scale production Slurm clusters with proper configuration (slurmctld, slurmdbd, partitions, QOS).
  • Connect customer clusters to the control plane and harmonize site schedulers and identities.
  • Provision bare metal hardware, manage firmware, rack networking, and out-of-band management.
  • Validate GPU nodes, manage driver stacks, NVLink, InfiniBand, and NCCL tuning.
  • Tune storage for high throughput and implement IaC pipelines (Ansible/Terraform).
  • Handle security escalation, STIG hardening, and Tier 3 incident response.
  • Lead on-call rotations and mentor junior engineers.

Skills

Linux system administration
Scripting (Bash, Python)
DevOps
Security hardening
Team collaboration

Education

Bachelor's degree in a related field

Tools

Slurm
Kubernetes/OpenShift
Ansible/Terraform
InfiniBand/NVLink networks

Job description

Parallel Works is seeking a Senior HPC Systems Engineer to design, operate and secure GPU-accelerated clusters behind defense and research programs. You will manage GPU node bring-up, Slurm configuration, fabric and storage troubleshooting, and security hardening across on-prem and government cloud environments.

The role is senior, handling escalations, coordinating with site staff, and mentoring junior engineers while ensuring high availability and compliance in risk-conscious environments.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior HPC Systems Engineer — Secure Hybrid GPU Clusters
Senior HPC Systems Engineer — Secure Hybrid GPU Clusters

Parallel Works • Chicago (IL)

Hybrid
USD 140,000 - 190,000
Medical, vision, dental coverage
401(k) with company match
Short term disability
+1
Senior HPC Systems Engineer — Slurm, GPU, Cloud-Native
Senior HPC Systems Engineer — Slurm, GPU, Cloud-Native

Nscale • New York (NY)

On-site
USD 180,000 - 260,000
Bonus
Equity
Medical Insurance
+5
Senior HPC & GPU Systems Engineer – TS/SCI Onsite
Senior HPC & GPU Systems Engineer – TS/SCI Onsite

VMD Corp • Bethesda (MD)

On-site
USD 140,000 - 180,000
Senior HPC & GPU Cluster Architect
Senior HPC & GPU Cluster Architect

sfcompute • San Francisco (CA)

On-site
USD 180,000 - 240,000
Generous equity grant
Visa Sponsorships
Retirement matching
+5
Senior HPC Systems Engineer
Senior HPC Systems Engineer

Parallel Works • United States

Hybrid
USD 180,000 - 230,000
Medical, vision, and dental coverage
401(k) with company match
Paid vacation & sick time
+1
Senior HPC Infrastructure Engineer
Senior HPC Infrastructure Engineer

Guardant Health • Palo Alto (CA)

Hybrid
USD 173,000 - 238,000
Hybrid work model
Senior HPC & GPU Cluster Architect — Scale & Automate
Senior HPC & GPU Cluster Architect — Scale & Automate

San Francisco Compute Company • San Francisco (CA)

Hybrid
USD 120,000 - 160,000
Generous equity grant
Competitive salary
Visa sponsorship
+6
Senior HPC & GPU Cluster Architect
Senior HPC & GPU Cluster Architect

The Consensus • San Francisco (CA)

Hybrid
USD 180,000 - 240,000
Visa sponsorships
401(k) retirement matching
Medical, dental & vision insurance
+2
Senior GPU/HPC Systems Engineer - Onsite Fremont
Senior GPU/HPC Systems Engineer - Onsite Fremont

Acceler8 Talent • Fremont (CA), Northern (KY)

Hybrid
USD 135,000 - 165,000
System Engineer
System Engineer

Acceler8 Talent • Fremont (CA), Northern (KY)

Hybrid
USD 135,000 - 165,000