Junior HPC Systems Engineer

Parallel Works

Deutschland

Vor Ort

EUR 56.000 - 86.000

Vollzeit

14 Tage+

Erhalte mehr Antworten von Arbeitgebern

Versende in nur wenigen Minuten einen passgenauen Lebenslauf.

Benefits dieser Stelle

Medical, vision, and dental coverage
401(k) with company match
Short term disability
Generous paid vacation and sick time

Zusammenfassung

Parallel Works is hiring a Junior HPC Systems Engineer to learn supercomputing operations on production systems. The role begins with monitoring, node health, and account/allocation work, and progresses to cluster builds and escalations.

The systems are hybrid: on-premises clusters, Government cloud regions, and commercial GPU providers. We seek engineers with solid Linux fundamentals and curiosity about how large systems behave; running a supercomputer is not a prerequisite.

Qualifikationen

  • 2+ years of hands-on Linux administration.
  • Understanding of HPC fundamentals: batch scheduling, shared filesystems, MPI.
  • Bash and Python scripting for operational work.
  • Practical networking: DNS, routing, firewalls, SSH keys, bastion access.
  • Git and a ticket-driven workflow; clear written communication.

Aufgaben

  • Monitor cluster health, node state, queue behavior, and alerting; respond to node failures and filesystem alerts.
  • Manage users, groups, and Slurm accounts; ensure consistency across venues.
  • Perform node lifecycle tasks: health checks, draining, and escalation for faults.
  • Extend Ansible playbooks and operational scripts; automate repetitive steps.
  • Patch and harden systems with senior review; provide evidence for security baselines.
  • Keep runbooks current and participate in on-call rotations after training.
  • Work with Linux distributions (RHEL, Rocky/Alma, Debian/Ubuntu) across environments.
  • Understand HPC concepts: MPI, shared storage, compute node health.
  • Collaborate on networking, VPNs, SSH access, and security practices.
  • Use Python/Bash scripting to automate operational tasks.
  • Exposure to production on-prem and cloud borders; interest in bare metal and site networks.
  • Possibility of U.S. citizenship and Secret clearance; sponsorship available for eligible candidates.

Kenntnisse

Linux administration
Bash scripting
Python scripting
Networking basics
Git
Clear written communication

Tools

Ansible
Terraform
nvidia-smi
DCGM

Jobbeschreibung

About Parallel Works

Parallel Works builds and operates ACTIVATE, a control plane for high performance computing and AI. Our customers run large scientific and AI workloads across their own on-premises clusters, Government and commercial cloud, and commercial GPU providers, and ACTIVATE gives them one way in to all of it. The high security boundary is authorized at Impact Level 5, with FIPS validated cryptography and STIG hardening throughout.


The work reaches most fields that depend on computing at scale: weather and climate forecasting, defense and intelligence programs, aerospace and structural analysis, molecular and materials science, energy, and AI research. A quarter here can include standing up a GPU cluster for one of those communities, federating a laboratory’s existing on-premises system with burst capacity it did not have before, and getting a domain code written decades ago to run on current hardware.


Customer success sets our priorities. We are a small engineering company, so engineers here work directly with the people using the systems and carry a problem from the first report through to the fix. This is what we call mission engineering: understanding what a customer is trying to accomplish and why the computing matters to it.


About the role

Parallel Works is hiring a Junior HPC Systems Engineer to learn supercomputing operations on production systems. The role starts with monitoring, node health, and account and allocation work, and moves into cluster builds and escalations.


The systems are hybrid: on-premises clusters the customer owns, accredited Government cloud regions, and commercial GPU providers, often in the same day. We are looking for strong Linux fundamentals and interest in how large systems behave. Running a supercomputer is not a prerequisite. The role leads into senior systems engineering, with cluster builds handled independently after about a year.


What you will do


  • Monitor and respond:cluster health, node state, queue behavior, and alerting, with first action on node failures, stuck jobs, and filesystem alerts.


  • Accounts and allocations:users, groups, Slurm accounts, and allocations, kept consistent across venues.


  • Node lifecycle:health checks on GPU and CPU nodes, draining and returning nodes, and the escalation path on faults, which is site staff or a vendor on customer hardware and a support case on cloud capacity.


  • Automation:extend Ansible playbooks and operational scripts, and replace the manual steps you find yourself repeating.


  • Patching and hardening:patches and hardening baselines under senior review, plus scan evidence for the security package.


  • Documentation and on call:keep runbooks current, and join the on call rotation once trained into it.


  • 2 or more years of hands-on Linux administration. RHEL, Rocky, or Alma, or Debian and Ubuntu; we run both families and you will work on both.


  • Understanding of HPC fundamentals: batch scheduling, shared filesystems, MPI job launch, and what makes a compute node healthy.


  • Bash and Python scripting for operational work.


  • Practical networking: DNS, routing, firewalls, SSH keys, bastion access.


  • Git and a ticket driven workflow, and clear written communication.


  • Interest in the on-premises side of the estate, meaning bare metal, out of band consoles, and site networking.


  • United States citizenship and eligibility for a Secret clearance, since the work reaches export controlled Government environments. An active clearance helps. We sponsor candidates who are eligible but not currently cleared.



Preferred Qualifications


  • Any exposure to Slurm, PBS Pro, or LSF.

  • Experience with Ansible, Terraform, or another configuration management tool.

  • Time in any public cloud, and with NVIDIA GPU tooling such as nvidia-smi or DCGM.

  • Time with physical servers: racking, BMC or IPMI consoles, firmware, or a home lab with hardware in it.

  • Involvement in a student cluster competition, a university HPC support role, or open source contributions.


Medical, vision, and dental coverage, a 401(k) with company match, short term disability, and generous paid vacation and sick time.


Equal employment opportunity

Parallel Works is an equal opportunity employer. We consider all qualified applicants without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, age, disability, protected veteran status, or any other characteristic protected by law.

Hol dir deinen kostenlosen, vertraulichen Lebenslauf-Check.
oder ziehe deine Datei hierhin.
Similar jobs

Ähnliche Jobs, die dir auch gefallen könnten

Senior HPC Systems Engineer
Senior HPC Systems Engineer

Parallel Works • Deutschland

Hybrid
EUR 90.000 - 130.000
Medical coverage
Vision coverage
Dental coverage
+3
HPC Systems Administrator
HPC Systems Administrator

EngineersOfAI • München

Vor Ort
EUR 60.000 - 80.000
Senior HPC Engineer, GPU Compute
Senior HPC Engineer, GPU Compute

United States Digital Space LLC • Berlin

Vor Ort
EUR 90.000 - 130.000
Competitive compensation
Career growth and learning
HPC Engineer Services
HPC Engineer Services

Embedded Shishya • Deutschland

Remote
EUR 60.000 - 80.000
System Engineer (Compute Node)
System Engineer (Compute Node)

United States Digital Space LLC • Berlin

Vor Ort
EUR 85.000 - 120.000
Competitive compensation
Career growth
Flexibility
+1
Talent Pool: HPC Architect
Talent Pool: HPC Architect

Quantori • Deutschland

Hybrid
EUR 70.000 - 100.000
Competitive compensation
Remote or office work
Healthcare benefits: medical insurance and paid sick leave
+2
Senior Infrastructure & Security Engineer
Senior Infrastructure & Security Engineer

European Tech Recruit • Frankfurt

Vor Ort
EUR 90.000 - 130.000
Senior System Engineer (Munich, Germany)
Senior System Engineer (Munich, Germany)

Remotestar • München

Hybrid
EUR 80.000 - 110.000
Indefinite contract
Equal pay guaranteed
Variable performance bonus
+8
Senior HPC Cluster Administrator - Deep Learning Frameworks Infrastructure
Senior HPC Cluster Administrator - Deep Learning Frameworks Infrastructure

NVIDIA • Berlin

Vor Ort
EUR 120.000 - 180.000
System Engineer
System Engineer

Aether Biomedical • Deutschland

Hybrid
EUR 70.000 - 110.000
Competitive compensation
Hybrid or remote options depending on役