Senior HPC Applications Engineer

Parallel Works, Inc.

Chicago (IL)

On-site

USD 140,000 - 210,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Medical, vision, and dental coverage
401(k) with company match
Paid vacation and sick time
Short term disability

Job summary

Parallel Works, Inc. is seeking a Senior HPC Applications Engineer to own the software stack and user experience on our ACTIVATE platform. You will support weather, AI, and scientific workloads across on-prem, cloud, and GPU providers, ensuring consistent behavior across compilers, modules, and filesystems.

The role involves escalation on build failures, multi-node runs, and throughput issues, with frequent collaboration with users to resolve root causes and deliver effective fixes.

Qualifications

  • 10+ years supporting scientific or AI application users on Linux HPC systems.
  • Building complex software from source on RHEL and Debian/Ubuntu systems: compilers, MPI, CMake and autotools.
  • Running Spack or EasyBuild and Lmod in production.
  • Experience with multi-node GPU workloads from the platform side.

Responsibilities

  • Software stack: compile and package MPI implementations (OpenMPI, MPICH, Intel MPI, HPC-X), compilers (GCC, Intel oneAPI, NVHPC), and scientific libraries via Spack or EasyBuild with Lmod module trees.
  • Enable AI and ML workloads: install requested frameworks, get multi-node GPU launch working, and diagnose environment vs framework issues.
  • Performance work: run scaling studies, profile with Nsight, VTune, TAU, HPCToolkit, or Score-P, and hand findings to the systems team.
  • Portability: get customer codes running on new GPU architectures and new venues using containers when needed.
  • User support: triage tickets, diagnose failed jobs, and provide written explanations.
  • Documentation and training: user guides, office hours, and training including Government sessions.

Skills

10+ years supporting scientific or AI
Building complex software from source
Spack or EasyBuild production
Lmod in production
Multi-node GPU workloads
Site-provided software stacks
Domain knowledge: weather/climate or C

Tools

Spack
EasyBuild
Lmod
MPI stacks (OpenMPI/MPICH/Intel MPI/HPC-X)

Job description

About Parallel Works

Parallel Works builds and operates ACTIVATE, a control plane for high performance computing and AI. Our customers run large scientific and AI workloads across their own on-premises clusters, Government and commercial cloud, and commercial GPU providers, and ACTIVATE gives them one way in to all of it. The high security boundary is authorized at Impact Level 5, with FIPS validated cryptography and STIG hardening throughout.

The work reaches most fields that depend on computing at scale: weather and climate forecasting, defense and intelligence programs, aerospace and structural analysis, molecular and materials science, energy, and AI research. A quarter here can include standing up a GPU cluster for one of those communities, federating a laboratory's existing on-premises system with burst capacity it did not have before, and getting a domain code written decades ago to run on current hardware.

Customer success sets our priorities. We are a small engineering company, so engineers here work directly with the people using the systems and carry a problem from the first report through to the fix. This is what we call mission engineering: understanding what a customer is trying to accomplish and why the computing matters to it.

About the role

Parallel Works is hiring a Senior HPC Applications Engineer to own the software stack and the user experience on our platforms. Our users are weather modelers, computational chemists, aerospace engineers, and AI researchers. The role is the escalation point for build failures, jobs that die partway through a multi-node run, and jobs running below expected throughput.

The scope covers both long-lived domain codes and current AI workloads, since customers run both on the same clusters. Users also move between on-premises systems, cloud, and commercial GPU providers, so a large part of the job is making an application behave the same across different compilers, site modules, MPI builds, and filesystems. Expect to spend a good share of the week talking to users.

What you will do
  • Software stack:compile and package MPI implementations (OpenMPI, MPICH, Intel MPI, HPC-X), compilers (GCC, Intel oneAPI, NVHPC), and scientific libraries, delivered through Spack or EasyBuild with Lmod module trees users can navigate.
  • Enable AI and ML workloads:install the frameworks customers ask for, get multi-node GPU launch working, and diagnose what sits below the framework: NCCL and collective behavior, container and driver mismatches, storage throughput, node faults mid-run. Customers drive their own toolchain choices.
  • Performance work:run scaling studies, profile with Nsight, VTune, TAU, HPCToolkit, or Score-P, and hand the finding to the systems team when the fix belongs in the fabric or the filesystem.
  • Portability:get customer codes running on new GPU architectures and new venues, using containers where that beats rebuilding against each site's modules.
  • User support:triage tickets, diagnose failed jobs to a root cause, and close them with a written explanation.
  • Documentation and training:user guides, office hours, and training for user communities, including formal Government training events.
  • 10 or more years supporting scientific or AI application users on Linux HPC systems.
  • Building complex software from source on both RHEL family and Debian or Ubuntu systems: compilers, MPI, CMake and autotools, and the dependency problems that come with them.
  • Running Spack or EasyBuild and Lmod in production.
  • Experience with multi-node GPU workloads from the platform side, and the judgment to tell a framework problem from an environment problem.
  • Working with site provided software stacks on on-premises systems as well as cloud images where you control the whole stack.
  • Working knowledge of at least one application domain: weather and climate, computational fluid dynamics, molecular and materials science, or structural analysis.
  • United States citizenship and eligibility for a Secret clearance, since the work reaches export controlled Government environments. An active clearance helps. We sponsor candidates who are eligible but not currently cleared.

You do not need every item on this list.

Preferred Qualifications
  • A prior user facing role at a Government supercomputing center, national laboratory, or university HPC center.
  • Depth in profiling and debugging tools: a GPU profiler, a CPU profiler, gdb, and MPI tooling.
  • Hands-on distributed training or inference work with PyTorch DDP or FSDP, DeepSpeed, Megatron style frameworks, JAX, vLLM, or TensorRT-LLM. Customers own their toolchains, so this is depth rather than a requirement.
  • Jupyter, remote visualization, or virtual desktop support for research users.
  • Medical, vision, and dental coverage, a 401(k) with company match, short term disability, and generous paid vacation and sick time.
Equal employment opportunity

Parallel Works is an equal opportunity employer. We consider all qualified applicants without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, age, disability, protected veteran status, or any other characteristic protected by law.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior HPC Applications Engineer
Senior HPC Applications Engineer

Parallel Works • Chicago (IL)

On-site
USD 140,000 - 190,000
Medical coverage
Vision coverage
Dental coverage
+2
Junior HPC Applications Engineer
Junior HPC Applications Engineer

Parallel Works • Chicago (IL)

On-site
USD 65,000 - 90,000
Medical coverage
Vision coverage
Dental coverage
+2
Junior HPC Systems Engineer
Junior HPC Systems Engineer

Parallel Works • Chicago (IL)

Hybrid
USD 70,000 - 100,000
Medical coverage
Vision coverage
Dental coverage
+4
Junior HPC Applications Engineer
Junior HPC Applications Engineer

Parallel Works, Inc. • Chicago (IL)

On-site
USD 65,000 - 90,000
Medical, vision, dental coverage
401(k) with company match
Short term disability
+1
Senior HPC Systems Engineer
Senior HPC Systems Engineer

Parallel Works • Chicago (IL)

Hybrid
USD 140,000 - 190,000
Medical, vision, dental coverage
401(k) with company match
Short term disability
+1
AI Compute Sales Lead
AI Compute Sales Lead

Parallel Works, Inc. • Chicago (IL)

On-site
USD 120,000 - 180,000
Sales commission
Health insurance
401(k) matching
+1
AI Compute Sales Lead
AI Compute Sales Lead

Parallel Works • Chicago (IL)

On-site
USD 120,000 - 180,000
Medical, vision, and dental coverage
401(k) with company match
Generous paid vacation and sick time
+1
HPC Software Engineer
HPC Software Engineer

Signature Federal Systems , LLC • Colorado Springs (CO)

On-site
USD 140,000 - 190,000
Defense Sales Lead
Defense Sales Lead

Parallel Works • Chicago (IL)

On-site
USD 120,000 - 190,000
Medical coverage
Vision coverage
Dental coverage
+4
Senior HPC Engineer for AI & Scientific Apps
Senior HPC Engineer for AI & Scientific Apps

Parallel Works • Chicago (IL)

On-site
USD 140,000 - 190,000
Medical coverage
Vision coverage
Dental coverage
+2