Staff HPC Infrastructure Architect (On-Prem & Cloud)

Guardant Health, Inc.

Palo Alto (CA)

On-site

USD 173,000 - 238,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Guardant Health, Inc. is seeking a Staff-level HPC Engineer to lead and scale our high-performance compute infrastructure across on-premise and cloud environments in Palo Alto, CA.

The role focuses on Linux systems, HPC file systems, and network/storage integration, with responsibilities spanning design, deployment, and 24/7 on-call support. You will collaborate with cross-functional teams to optimize performance, reliability, and automation using Slurm, Docker, Kubernetes, and Warewulf, while

Qualifications

  • Bachelor’s degree in Computer Science or a related field with 8–12 years of relevant experience; Master’s degree with 6–8 years of relevant experience; or PhD with 3–5 years of relevant experience.
  • Strong experience in systems and/or infrastructure engineering, including Linux/Unix administration and TCP/IP networking
  • Hands-on experience with automation tools, such as Ansible or equivalent technologies
  • Experience with high-performance networking technologies, such as InfiniBand, RoCE, RDMA, or equivalent, including troubleshooting in production environments
  • Experience supporting large-scale data storage and high-performance computing (HPC)/compute environments
  • Experience working with both on-premise and cloud-based infrastructure, such as AWS, Google Cloud Platform (GCP), Azure, or similar environments
  • Experience developing and supporting software release, operations, and infrastructure automation processes and toolsets
  • Strong experience creating and maintaining system administration and technical documentation

Responsibilities

  • Manage multiple HPC clusters and cluster file systems
  • Integrate cloud bursting as part of the HPC abstraction work
  • Research, develop, and implement the next generation HPC solutions
  • Troubleshoot the production system stack down to source code level, e.g shell scripts, Python, and others
  • Maintain, monitor, and support the infrastructure environment and/or facilities
  • Use and maintain enhanced production monitoring and addition capability
  • Support improvements for increased system reliability and performance
  • Support multiple systems or applications of medium to high complexity defined by size, technology used, and system feeds and interfaces) with multiple concurrent users, ensuring control, integrity, and accessibility
  • Support systems at remote locations, including internationally
  • Mentor junior engineers on HPC best practices
  • Work with offsite consultants to maintain the infrastructure
  • Work with vendors to troubleshoot, upgrade, and repair systems as needed
  • Represent HPC infrastructure networking and storage-integration topics in cross-functional planning with networking, SQA, DevOps/SRE, and the MSP
  • Set up and ownership supporting XDMoD instances for HPC metric and monitoring
  • Participate in a 24/7 on-call rotation
  • Act as the technical peer for HPC networking and interconnect initiatives with the dedicated networking engineer
  • Collaborate on design, performance tuning, and troubleshooting of HPC Ethernet
  • Work with enterprise networking on integration of HPC systems with the bandwidth-on-demand system that connects our sites and cloud infrastructure
  • Work with the networking infrastructure team to manage and optimize connectivity to and from HPC systems and global locations
  • Act as the technical peer for the architecture and integration strategy for. HPC storage in partnership with the dedicated storage engineer and MSP
  • Serve as a technical point of contact for the MSP storage relationship and help define and evolve SLAs, validate delivery, and elevate technical issues
  • Support the transition of day-to-day storage operations to the MSP without loss of performance or reliability
  • Strong experience with cloud bursting technologies
  • Experience with wide area file systems
  • Experience with Docker and Apptainer container technologies
  • Experience with Kubernetes
  • Operating infrastructure compliant with HIPAA and SOX standards

Skills

Linux/Unix administration
TCP/IP networking
Ansible
Slurm
Kubernetes
Docker
Cloud platforms (AWS/GCP/Azure)

Education

Bachelor’s degree in Computer Science or related field

Tools

GPFS
Warewulf
InfiniBand/RDMA

Job description

Guardant Health, Inc. is seeking a Staff-level HPC Engineer to lead and scale our high-performance compute infrastructure across on-premise and cloud environments in Palo Alto, CA.

The role focuses on Linux systems, HPC file systems, and network/storage integration, with responsibilities spanning design, deployment, and 24/7 on-call support. You will collaborate with cross-functional teams to optimize performance, reliability, and automation using Slurm, Docker, Kubernetes, and Warewulf, while

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior HPC Infrastructure Engineer
Senior HPC Infrastructure Engineer

Guardant Health • Palo Alto (CA)

Hybrid
USD 173,000 - 238,000
Hybrid work model
Staff HPC Infrastructure Engineer
Staff HPC Infrastructure Engineer

Guardant Health, Inc. • Palo Alto (CA)

On-site
USD 173,000 - 238,000
Senior HPC Systems Engineer — Secure Hybrid GPU Clusters
Senior HPC Systems Engineer — Secure Hybrid GPU Clusters

Parallel Works • Chicago (IL)

Hybrid
USD 140,000 - 190,000
Medical, vision, dental coverage
401(k) with company match
Short term disability
+1
Senior HPC Linux Systems Engineer (Storage &DataProtection)
Senior HPC Linux Systems Engineer (Storage &DataProtection)

10x Genomics Inc • Pleasanton (CA)

Hybrid
USD 137,000 - 185,000
Equity grants
Health benefits
Retirement plan
+1
Senior HPC Architect – On-Prem/Cloud/Hybrid Computing
Senior HPC Architect – On-Prem/Cloud/Hybrid Computing

BioSpace • North Chicago (IL)

On-site
USD 150,000 - 210,000
Paid time off
Medical/dental/vision insurance
401(k)
+1
Senior HPC & GPU Cluster Architect
Senior HPC & GPU Cluster Architect

The Consensus • San Francisco (CA)

Hybrid
USD 180,000 - 240,000
Visa sponsorships
401(k) retirement matching
Medical, dental & vision insurance
+2
HPC Architect: Cloud-Driven Scientific Compute Leader
HPC Architect: Cloud-Driven Scientific Compute Leader

Allergan • Chicago (IL), Northern (KY)

Hybrid
USD 180,000 - 240,000
HPC Systems Engineer — Remote/Hybrid, Slurm/Linux
HPC Systems Engineer — Remote/Hybrid, Slurm/Linux

Strategic Business Systems, Inc (SBS) • Chantilly (VA)

Hybrid
USD 120,000 - 180,000
Flexible work arrangements
Senior HPC Platform Engineer — Hybrid Linux & Cloud
Senior HPC Platform Engineer — Hybrid Linux & Cloud

X-ISS • United States

Hybrid
USD 100,000 - 130,000
401k matching
Health insurance
Vision insurance
+3
Senior HPC & GPU Cluster Architect
Senior HPC & GPU Cluster Architect

sfcompute • San Francisco (CA)

On-site
USD 180,000 - 240,000
Generous equity grant
Visa Sponsorships
Retirement matching
+5