Senior HPC Systems Engineer

Parallel Works

United States

Hybrid

USD 180,000 - 230,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Medical, vision, and dental coverage
401(k) with company match
Paid vacation & sick time
Short term disability

Job summary

Parallel Works is seeking a Senior HPC Systems Engineer to design, operate and secure GPU-accelerated clusters behind defense and research programs. You will manage GPU node bring-up, Slurm configuration, fabric and storage troubleshooting, and security hardening across on-prem and government cloud environments.

The role is senior, handling escalations, coordinating with site staff, and mentoring junior engineers while ensuring high availability and compliance in risk-conscious environments.

Qualifications

  • 10+ years operating production Linux systems across multiple distributions (RHEL/ Rocky/ Alma on Gov side; Debian/Ubuntu on commercial GPU side).
  • Experience with Slurm, HPC scheduling, and cluster hardening in a hybrid environment.
  • Proficiency with Bash and Python scripting.
  • DevOps tooling experience with infrastructure as code (Ansible, Terraform).
  • Security and incident response involving STIGs, FIPS cryptography, and secure network configurations.

Responsibilities

  • Operate and scale production Slurm clusters with proper configuration (slurmctld, slurmdbd, partitions, QOS).
  • Connect customer clusters to the control plane and harmonize site schedulers and identities.
  • Provision bare metal hardware, manage firmware, rack networking, and out-of-band management.
  • Validate GPU nodes, manage driver stacks, NVLink, InfiniBand, and NCCL tuning.
  • Tune storage for high throughput and implement IaC pipelines (Ansible/Terraform).
  • Handle security escalation, STIG hardening, and Tier 3 incident response.
  • Lead on-call rotations and mentor junior engineers.

Skills

Linux system administration
Scripting (Bash, Python)
DevOps
Security hardening
Team collaboration

Education

Bachelor's degree in a related field

Tools

Slurm
Kubernetes/OpenShift
Ansible/Terraform
InfiniBand/NVLink networks

Job description

About Parallel Works

Parallel Works builds and operates ACTIVATE, a control plane for high performance computing and AI. Our customers run large scientific and AI workloads across their own on-premises clusters, Government and commercial cloud, and commercial GPU providers, and ACTIVATE gives them one way in to all of it. The high security boundary is authorized at Impact Level 5, with FIPS validated cryptography and STIG hardening throughout.

The work reaches most fields that depend on computing at scale: weather and climate forecasting, defense and intelligence programs, aerospace and structural analysis, molecular and materials science, energy, and AI research. A quarter here can include standing up a GPU cluster for one of those communities, federating a laboratory’s existing on-premises system with burst capacity it did not have before, and getting a domain code written decades ago to run on current hardware.

Customer success sets our priorities. We are a small engineering company, so engineers here work directly with the people using the systems and carry a problem from the first report through to the fix. This is what we call mission engineering: understanding what a customer is trying to accomplish and why the computing matters to it.

About the role

Parallel Works is hiring a Senior HPC Systems Engineer to build and run the clusters behind our defense and research programs. The work covers GPU node bring-up, Slurm configuration, fabric and storage troubleshooting, security hardening, and Tier 3 escalation.

The computing environments are hybrid. Some clusters are customer owned hardware on site, some run in accredited Government cloud regions, and some are dedicated GPU clusters at commercial providers. On several programs the on-premises systems carry the primary load and cloud takes the overflow. The position is senior: it handles the escalations the rest of the team cannot resolve, and it trains the junior engineers.

What you will do
  • Cluster operations:build and operate production Slurm clusters. slurmctld and slurmdbd, partitions and QOS, accounts and fair share, GPU GRES, prolog and epilog, cgroup enforcement.
  • Hybrid federation:connect customer owned clusters to the control plane, reconciling their site scheduler, storage, and identity source so accounts and allocations behave the same in every venue.
  • On-premises hardware:bare metal provisioning, out of band management, firmware, rack networking, and fault coordination with site staff or vendors.
  • GPU and fabric:validate GPU nodes before users arrive. Driver and CUDA stack, DCGM health checks, XID triage, fabric manager and NVLink checks, InfiniBand verification, NCCL tuning.
  • Storage and automation:tune parallel and high throughput storage, and write the Ansible, Terraform, and image build pipelines that make a cluster reproducible.
  • Security and escalation:STIG hardening, scan remediation, FIPS validated cryptography, security package artifacts, Tier 3 escalations, and a share of the on call rotation.
  • 10 or more years operating production Linux systems across more than one distribution family. RHEL, Rocky, or Alma on the Government side and Debian or Ubuntu on the commercial GPU side, since those clusters usually ship Ubuntu. Kernel and network tuning, systemd, cgroups, NUMA.
  • Production Slurm administration. You have configured, debugged, and upgraded a scheduler other people depended on.
  • At least one parallel or high throughput filesystem in production, plus InfiniBand or RoCE fabric operations.
  • NVIDIA GPU node operations at multi-node scale, including driver stack management and fault triage.
  • Experience on customer owned or on-premises clusters as well as public cloud, including bare metal provisioning and out of band management.
  • DevOps experience with infrastructure as code frameworks such as Ansible or Terraform.
  • Proficiency with common scripting languages such as Bash and Python.
  • United States citizenship and eligibility for a Secret clearance, since the work reaches export controlled Government environments. An active clearance helps. We sponsor candidates who are eligible but not currently cleared.

You do not need every item on this list. If you have most of it and work well with other people, apply.

Preferred Qualifications
  • Time at a Government supercomputing center, national laboratory, or university research computing center.
  • Work with STIG, SCAP, Tenable, eMASS, or RMF, or time inside FedRAMP or Impact Level boundaries.
  • Running GPU workloads on Kubernetes or OpenShift with Helm and operators, and using HPC containers such as Apptainer, Enroot, or Pyxis.
  • Running PBS Pro alongside Slurm, and monitoring with Prometheus and Grafana, including utilization and chargeback reporting.

Medical, vision, and dental coverage, a 401(k) with company match, short term disability, and generous paid vacation and sick time.

Equal employment opportunity

Parallel Works is an equal opportunity employer. We consider all qualified applicants without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, age, disability, protected veteran status, or any other characteristic protected by law.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Junior HPC Systems Engineer
Junior HPC Systems Engineer

Parallel Works • Chicago (IL)

On-site
USD 70,000 - 100,000
Medical coverage
Vision coverage
Dental coverage
+3
Senior HPC Applications Engineer
Senior HPC Applications Engineer

Parallel Works, Inc. • Chicago (IL)

On-site
USD 140,000 - 210,000
Medical, vision, and dental coverage
401(k) with company match
Paid vacation and sick time
+1
Junior HPC Applications Engineer
Junior HPC Applications Engineer

Parallel Works, Inc. • Chicago (IL)

On-site
USD 65,000 - 90,000
Medical, vision, dental coverage
401(k) with company match
Short term disability
+1
AI Compute Sales Lead
AI Compute Sales Lead

Parallel Works, Inc. • Chicago (IL)

On-site
USD 120,000 - 180,000
Sales commission
Health insurance
401(k) matching
+1
Senior HPC Systems Engineer: Slurm, GPU & Hybrid Clusters
Senior HPC Systems Engineer: Slurm, GPU & Hybrid Clusters

Parallel Works • United States

Hybrid
USD 180,000 - 230,000
Medical, vision, and dental coverage
401(k) with company match
Paid vacation & sick time
+1
Systems Engineer – HPC & GPU Infrastructure
Systems Engineer – HPC & GPU Infrastructure

FiveInsights • Bethesda (MD)

On-site
USD 170,000 - 210,000
Systems Engineer – HPC & GPU Infrastructure
Systems Engineer – HPC & GPU Infrastructure

MAXISIQ, Inc. • Bethesda (MD)

On-site
USD 170,000 - 210,000
Member of Technical Staff - AI Infrastructure
Member of Technical Staff - AI Infrastructure

Veeda Innovation • California (MO)

Hybrid
USD 180,000 - 240,000
Senior HPC Systems Engineer — Secure Hybrid GPU Clusters
Senior HPC Systems Engineer — Secure Hybrid GPU Clusters

Parallel Works • Chicago (IL)

Hybrid
USD 140,000 - 190,000
Medical, vision, dental coverage
401(k) with company match
Short term disability
+1
HPC Infrastructure Engineer
HPC Infrastructure Engineer

Arcadia • San Francisco (CA)

On-site
USD 180,000 - 260,000