Senior HPC Systems Engineer — Secure Hybrid GPU Clusters

Parallel Works

Chicago (IL)

Hybrid

USD 140,000 - 190,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Medical, vision, dental coverage
401(k) with company match
Short term disability
Generous paid vacation and sick time

Job summary

Parallel Works seeks a Senior HPC Systems Engineer to build and run production Slurm clusters and oversee GPU nodes across on-prem and cloud environments. You will handle bare metal provisioning, network and firmware tasks, and ensure secure, scalable operation for defense and research programs.

The role is senior and involves triage escalations, mentoring junior engineers, and contributing to automation with Ansible or Terraform. U.S.

Qualifications

  • 10+ years operating production Linux systems across multiple distributions.
  • Experience with SLURM and high-throughput HPC environments.
  • Experience with on-prem hardware, GPU clusters, and fabric networks.
  • Proficiency in Bash and Python scripting for automation.
  • Familiarity with DevOps IaC tools (Ansible, Terraform).
  • Security hardening and compliance in government or enterprise settings.

Responsibilities

  • Cluster operations: build and run production Slurm clusters with proper accounting and QoS.
  • Hybrid federation: connect customer clusters to the control plane and align site schedulers.
  • On-prem hardware: bare metal provisioning, out-of-band management, firmware, networking.
  • GPU and fabric: validate GPU nodes, driver stacks, NCCL tuning, InfiniBand checks.
  • Storage and automation: optimize storage throughput and develop reproducible pipelines.
  • Security and escalation: STIG hardening, cryptography, Tier 3 escalations.

Skills

Production Linux systems
RHEL Rocky Alma
Debian Ubuntu Linux
kernel tuning
systemd
cgroups
NUMA
Slurm administration
InfiniBand fabric
NVIDIA GPU management
Ansible Terraform
Bash Python scripting

Tools

Prometheus
Grafana

Job description

Parallel Works seeks a Senior HPC Systems Engineer to build and run production Slurm clusters and oversee GPU nodes across on-prem and cloud environments. You will handle bare metal provisioning, network and firmware tasks, and ensure secure, scalable operation for defense and research programs.

The role is senior and involves triage escalations, mentoring junior engineers, and contributing to automation with Ansible or Terraform. U.S.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior HPC & GPU Systems Engineer
Senior HPC & GPU Systems Engineer

Xcelerate-Solutions-5 • Bethesda (MD)

On-site
USD 120,000 - 160,000
Senior HPC Systems Engineer
Senior HPC Systems Engineer

Parallel Works • Chicago (IL)

On-site
USD 140,000 - 190,000
Medical, vision, dental coverage
401(k) with company match
Short term disability
+1
Senior HPC Systems Engineer — Slurm, GPU, Cloud-Native
Senior HPC Systems Engineer — Slurm, GPU, Cloud-Native

Nscale • New York (NY)

On-site
USD 180,000 - 260,000
Bonus
Equity
Medical Insurance
+5
Senior Systems Engineer – HPC & GPU Infrastructure
Senior Systems Engineer – HPC & GPU Infrastructure

Xcelerate Solutions • Bethesda (MD)

On-site
USD 130,000 - 180,000
Senior HPC & GPU Cluster Architect — Scale & Automate
Senior HPC & GPU Cluster Architect — Scale & Automate

San Francisco Compute Company • San Francisco (CA)

Hybrid
USD 120,000 - 160,000
Generous equity grant
Competitive salary
Visa sponsorship
+6
Senior HPC Support Engineer — Linux, GPUs, Kubernetes
Senior HPC Support Engineer — Linux, GPUs, Kubernetes

Lambda Inc. • United States

On-site
USD 150,000 - 190,000
Health, dental, and vision coverage
401k Plan with 2% company match
Wellness and commuter stipends
+1
Senior HPC Support Engineer - GPU Cloud Infra
Senior HPC Support Engineer - GPU Cloud Infra

Neura Market • United States

On-site
USD 150,000 - 190,000
Wellness stipend
Commuter stipend
401k plan with 2% company match (USA)
+1
HPC AI Systems Architect (On-Prem GPU Cluster)
HPC AI Systems Architect (On-Prem GPU Cluster)

MRE Consulting • Houston (TX)

On-site
USD 95,000 - 140,000
Senior HPC Systems Engineer for High-Performance Clusters
Senior HPC Systems Engineer for High-Performance Clusters

SPACE EXPLORATION TECHNOLOGIES CORP • United States

On-site
USD 140,000 - 210,000
Junior HPC Applications Engineer — Launch Research Clusters
Junior HPC Applications Engineer — Launch Research Clusters

Parallel Works, Inc. • Chicago (IL)

On-site
USD 65,000 - 90,000
Medical, vision, dental coverage
401(k) with company match
Short term disability
+1