Senior HPC Infrastructure Engineer: Clusters & Cloud

Jobtailor

California (MO)

On-site

USD 150,000 - 210,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Jobtailor seeks an experienced HPC Infrastructure Engineer to manage multiple HPC clusters, integrate cloud bursting, and develop next‑gen HPC solutions across on‑prem and cloud environments. You will troubleshoot production stacks to source‑code level, mentor junior engineers, and participate in a 24/7 on‑call rotation while collaborating with MSP and vendors to ensure reliability and performance.

Expertise in Linux/TCP/IP, automation with Ansible, and HPC storage—plus familiarity with

Qualifications

  • Bachelor’s degree in Computer Science or related field with 8–12 years of experience; Master’s with 6–8 years; PhD with 3–5 years.
  • Strong Linux/Unix and TCP/IP networking experience.
  • Hands-on automation with Ansible or equivalent.
  • Experience with InfiniBand, RoCE, RDMA, or equivalent.
  • Experience with on-prem and cloud infra (AWS, GCP, Azure).
  • Experience developing release, operations, and infrastructure automation processes.

Responsibilities

  • Manage multiple HPC clusters and cluster file systems.
  • Integrate cloud bursting into HPC abstraction work.
  • Research, develop, and implement next-generation HPC solutions.
  • Troubleshoot production stack to source-code level (shell scripts, Python).
  • Maintain, monitor, and support infrastructure environments and facilities.
  • Improve production monitoring capabilities.

Skills

HPC Cluster Management
Linux/Unix Administration
High-Performance Networking
Automation (Ansible)
Cloud Infrastructure (AWS, GCP, Azure)

Education

Bachelor’s degree in CS or related field
Master’s degree (advanced)

Tools

Ansible
Docker
Kubernetes
Slurm
Warewulf
GPFS
Red Hat Linux

Job description

Jobtailor seeks an experienced HPC Infrastructure Engineer to manage multiple HPC clusters, integrate cloud bursting, and develop next‑gen HPC solutions across on‑prem and cloud environments. You will troubleshoot production stacks to source‑code level, mentor junior engineers, and participate in a 24/7 on‑call rotation while collaborating with MSP and vendors to ensure reliability and performance.

Expertise in Linux/TCP/IP, automation with Ansible, and HPC storage—plus familiarity with

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior HPC Platform Engineer — Hybrid Linux & Cloud
Senior HPC Platform Engineer — Hybrid Linux & Cloud

X-ISS • United States

Hybrid
USD 100,000 - 130,000
401k matching
Health insurance
Vision insurance
+3
Senior HPC Systems Engineer — Secure Hybrid GPU Clusters
Senior HPC Systems Engineer — Secure Hybrid GPU Clusters

Parallel Works • Chicago (IL)

Hybrid
USD 140,000 - 190,000
Medical, vision, dental coverage
401(k) with company match
Short term disability
+1
Senior HPC Infrastructure Engineer – Remote
Senior HPC Infrastructure Engineer – Remote

Guardant Health, Inc. • United States

Hybrid
USD 156,000 - 214,000
HPC Cluster Engineer: Secure, High-Performance Compute
HPC Cluster Engineer: Secure, High-Performance Compute

D2 Technical Services • Springfield (VA)

On-site
USD 170,000 - 180,000
Health/Dental/Vision
401(k) match
PTO
Lead HPC Cluster Engineer for AI/ML & OpenShift
Lead HPC Cluster Engineer for AI/ML & OpenShift

Abile Group, Inc • Springfield (VA)

On-site
USD 130,000 - 180,000
Senior Systems Engineer: HPC & Cloud (Hybrid)
Senior Systems Engineer: HPC & Cloud (Hybrid)

Holtec International • Camden (NJ)

On-site
USD 135,000 - 155,000
Medical, dental, and vision coverage
Hybrid work opportunities
401(k) with company match
+5
Senior HPC & GPU Cluster Architect
Senior HPC & GPU Cluster Architect

The Consensus • San Francisco (CA)

Hybrid
USD 180,000 - 240,000
Visa sponsorships
401(k) retirement matching
Medical, dental & vision insurance
+2
Senior HPC Systems Engineer: Scale AI Clusters
Senior HPC Systems Engineer: Scale AI Clusters

SpaceX • Hawthorne (CA)

On-site
USD 165,000 - 230,000
Stock options
Health, vision and dental coverage
401(k) retirement plan
+2
Senior HPC & AI Cloud Solutions Architect
Senior HPC & AI Cloud Solutions Architect

Career Techniques • Dallas (TX)

Hybrid
USD 110,000 - 145,000
HPC Cluster Engineer — AI/ML & OpenShift Infra
HPC Cluster Engineer — AI/ML & OpenShift Infra

Linuxconfig • Springfield (VA)

Hybrid
USD 140,000 - 185,000