Junior HPC Engineer: Build, Monitor, and Scale Clusters

Parallel Works

Chicago (IL)

Hybrid

USD 70,000 - 100,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Medical coverage
Vision coverage
Dental coverage
401(k) with company match
Paid vacation
Sick time

Job summary

Parallel Works seeks a Junior HPC Systems Engineer to learn supercomputing operations on production systems. The role starts with monitoring, node health, and account work, then moves into cluster builds and escalations across hybrid on-prem and cloud environments.

Strong Linux fundamentals are required; experience with Slurm and HPC concepts is helpful. The position offers growth into senior roles with hands-on responsibility.

Qualifications

  • 2+ years of hands-on Linux administration across RHEL, Rocky, Alma, Debian and Ubuntu.
  • HPC fundamentals: batch scheduling, shared filesystems, MPI launch.
  • Bash and Python scripting for operational tasks.
  • Practical networking: DNS, routing, firewalls, SSH keys, bastion access.
  • Git and a ticket-driven workflow with clear written communication.
  • Interest in on-premises estate: bare metal, out-of-band consoles, site networking.
  • US citizenship and eligibility for a Secret clearance; active clearance helpful but not required; we sponsor eligible candidates.

Responsibilities

  • Monitor cluster health, node state, queue behavior and alerting; act on node failures and filesystem alerts.
  • Manage users, groups, Slurm accounts and allocations consistently across venues.
  • Oversee node lifecycle: health checks, draining and returning nodes, escalate faults.
  • Extend Ansible playbooks and scripts to automate repetitive tasks.
  • Patch and harden systems; provide security baselines and evidence for audits.
  • Keep runbooks current and participate in on-call rotation after training.

Skills

Linux administration
Bash scripting
Python scripting
Networking
Git and ticket workflow
On-premises infrastructure
Cloud familiarity
NVIDIA GPU tooling
Student cluster/open source

Tools

Ansible
Terraform
Git
nvidia-smi/DCGM
BMC/IPMI

Job description

Parallel Works seeks a Junior HPC Systems Engineer to learn supercomputing operations on production systems. The role starts with monitoring, node health, and account work, then moves into cluster builds and escalations across hybrid on-prem and cloud environments.

Strong Linux fundamentals are required; experience with Slurm and HPC concepts is helpful. The position offers growth into senior roles with hands-on responsibility.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Junior HPC Applications Engineer — Launch Research Clusters
Junior HPC Applications Engineer — Launch Research Clusters

Parallel Works, Inc. • Chicago (IL)

On-site
USD 65,000 - 90,000
Medical, vision, dental coverage
401(k) with company match
Short term disability
+1
Junior HPC Systems Engineer
Junior HPC Systems Engineer

Parallel Works • Chicago (IL)

On-site
USD 70,000 - 100,000
Medical coverage
Vision coverage
Dental coverage
+3
Senior HPC Systems Engineer — Secure Hybrid GPU Clusters
Senior HPC Systems Engineer — Secure Hybrid GPU Clusters

Parallel Works • Chicago (IL)

Hybrid
USD 140,000 - 190,000
Medical, vision, dental coverage
401(k) with company match
Short term disability
+1
Junior HPC Applications Engineer
Junior HPC Applications Engineer

Parallel Works, Inc. • Chicago (IL)

On-site
USD 65,000 - 90,000
Medical, vision, dental coverage
401(k) with company match
Short term disability
+1
Senior HPC Systems Engineer – Clusters, Linux & AI
Senior HPC Systems Engineer – Clusters, Linux & AI

United States Digital Space LLC • El Segundo (CA)

On-site
USD 165,000 - 265,000
Medical, Vision & Dental
401(k) plan
Disability insurance
+4
Systems Administration - HPC Cluster
Systems Administration - HPC Cluster

Metasys Technologies • Newton (MA)

On-site
USD 96,432 - 99,187
Senior HPC Systems Engineer: Scale AI Clusters
Senior HPC Systems Engineer: Scale AI Clusters

SpaceX • Hawthorne (CA)

On-site
USD 165,000 - 230,000
Stock options
Health, vision and dental coverage
401(k) retirement plan
+2
Senior HPC Infrastructure Engineer
Senior HPC Infrastructure Engineer

Guardant Health • Palo Alto (CA)

Hybrid
USD 173,000 - 238,000
Hybrid work model
Senior System Engineer, HPC/AI & Cloud Clusters
Senior System Engineer, HPC/AI & Cloud Clusters

Supermicro • San Jose (CA)

On-site
USD 137,000 - 156,000
HPC Systems Engineer — Remote/Hybrid, Slurm/Linux
HPC Systems Engineer — Remote/Hybrid, Slurm/Linux

Strategic Business Systems, Inc (SBS) • Chantilly (VA)

Hybrid
USD 120,000 - 180,000
Flexible work arrangements