Lead GPU HPC Infrastructure Engineer

Sciforium

San Francisco (CA)

On-site

USD 180,000 - 240,000

Full time

3 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Medical, dental, and vision insurance
401k plan
Daily lunch, snacks, and beverages
Flexible time off
Competitive salary and equity

Job summary

Sciforium is seeking a Senior HPC & GPU Infrastructure Engineer to own the health, reliability, and performance of a high-density GPU compute cluster. You will bridge hardware operations, distributed systems, and ML workflows, from Linux systems engineering to CUDA/ROCm stacks and vLLM debugging.

Ideal candidates have 5+ years in HPC/GPU environments, with deep Linux internals knowledge and strong scripting (Bash/Python).

Qualifications

  • Requires 5+ years in HPC, GPU cluster operations or similar roles.
  • Bachelor’s or Master’s in a technical field.
  • Strong CUDA/ROCm and GPU driver debugging experience.

Responsibilities

  • Own health, reliability, and performance of the GPU compute cluster.
  • Lead deployment of new GPU nodes and topology validation.
  • Maintain ML software stack (CUDA, PyTorch, JAX, vLLM) and driver stacks.

Skills

Linux systems engineering
SRE / reliability engineering
Kernel debugging
Networking security
Automation scripting
Python scripting

Education

Bachelor’s or Master’s degree in CS/CE/EE or related field

Tools

NVIDIA GPUs (H100/B200)
AMD GPUs (MI325x/MI355x)
CUDA Toolkit / cuDNN / NCCL
ROCm / ROCm stack
Linux kernel modules
GPFS / Lustre / NFS
NDMA / RDMA networking
SSH / VPN / iptables

Job description

Sciforium is seeking a Senior HPC & GPU Infrastructure Engineer to own the health, reliability, and performance of a high-density GPU compute cluster. You will bridge hardware operations, distributed systems, and ML workflows, from Linux systems engineering to CUDA/ROCm stacks and vLLM debugging.

Ideal candidates have 5+ years in HPC/GPU environments, with deep Linux internals knowledge and strong scripting (Bash/Python).

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior GPU Cluster Engineer for AI Infrastructure
Senior GPU Cluster Engineer for AI Infrastructure

Sciforium • San Francisco (CA)

On-site
USD 150,000 - 220,000
Medical, dental, and vision insurance
401k plan
Daily lunch, snacks, and beverages
+2
Senior HPC & GPU Infrastructure Engineer
Senior HPC & GPU Infrastructure Engineer

Sciforium • San Francisco (CA)

On-site
USD 180,000 - 240,000
Medical, dental, and vision insurance
401k plan
Daily lunch, snacks, and beverages
+2
GPU Cluster Engineer, Systems & Platform
GPU Cluster Engineer, Systems & Platform

Sciforium • San Francisco (CA)

On-site
USD 150,000 - 220,000
Medical, dental, and vision insurance
401k plan
Daily lunch, snacks, and beverages
+2
GPU Cluster Architect: Scalable AI Platform
GPU Cluster Architect: Scalable AI Platform

Sciforium • San Francisco (CA)

On-site
USD 190,000 - 270,000
Medical insurance
401k plan
Daily meals/snacks
+2
GPU Cluster Engineer, Systems & Platform
GPU Cluster Engineer, Systems & Platform

Sciforium • San Francisco (CA)

On-site
USD 190,000 - 270,000
Medical insurance
401k plan
Daily meals/snacks
+2
Senior GPU Compute Cluster Engineer
Senior GPU Compute Cluster Engineer

Inferact Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 200,000 - 400,000
Health benefits
Dental benefits
Vision benefits
+1
Senior GPU Cluster Networking Engineer (RDMA/InfiniBand)
Senior GPU Cluster Networking Engineer (RDMA/InfiniBand)

Sciforium • San Francisco (CA)

On-site
USD 170,000 - 230,000
Remote GPU Cluster Engineer & Automation Lead
Remote GPU Cluster Engineer & Automation Lead

AI Chopping Block • Northern (KY)

Hybrid
USD 150,000 - 230,000
Senior GPU Infra Lead: Slurm, Kubernetes & Platform
Senior GPU Infra Lead: Slurm, Kubernetes & Platform

Jobgether SRL • United States

Remote
USD 170,000 - 250,000
HPC Infrastructure Engineer
HPC Infrastructure Engineer

Arcadia • San Francisco (CA)

On-site
USD 180,000 - 260,000