Senior GPU Infra Architect with Equity

Sciforium

San Francisco (CA)

On-site

USD 150,000 - 220,000

Full time

14 hours ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Medical, dental, and vision
401k plan
Lunch and beverages
Flexible time off
Salary and equity

Job summary

Sciforium in San Francisco is seeking a Senior HPC & GPU Infrastructure Engineer to own the health, reliability, and performance of our GPU compute cluster. You will be the primary custodian of a high-density accelerator environment, bridging hardware operations, distributed systems, and ML workflows.

This role involves hands-on Linux systems engineering, GPU driver bring-up, and maintaining the ML software stack (CUDA/ROCm, PyTorch, JAX, vLLM).

Qualifications

  • 5+ years in HPC, GPU cluster ops, or related Linux roles.
  • BS/MS in CS/CE/EE or related field.
  • Deep knowledge of NVIDIA or AMD GPUs, drivers and kernel debugging.

Responsibilities

  • Ensure system health, reliability, and performance of the GPU compute cluster.
  • Lead on-call responses for outages, GPU failures, and node crashes.
  • Develop monitoring for GPU health, memory errors, and topology issues.
  • Coordinate vendor and data-center activities for repairs and RMAs.
  • Manage Linux OS, patching, kernel tuning, and automation for large fleets.
  • Secure infrastructure with VPNs, firewalls, SSH hardening, and access controls.
  • Deploy and bring up new GPU nodes, BIOS, NUMA tuning, and topology validation.
  • Maintain ML stacks: PyTorch, JAX, CUDA toolkit, cuDNN, ROCm, NCCL.

Skills

Bash scripting
Python scripting
Linux internals
GPU driver debugging
CUDA toolkit
ROCm stack
Networking security
RDMA networking

Education

Bachelor's or Master's degree in Computer Science/Computer Engineering/Electrical Engineering

Tools

Slurm
Kubernetes
Run:AI
Ansible
SaltStack
Terraform
NVIDIA drivers
ROCm tooling

Job description

Sciforium in San Francisco is seeking a Senior HPC & GPU Infrastructure Engineer to own the health, reliability, and performance of our GPU compute cluster. You will be the primary custodian of a high-density accelerator environment, bridging hardware operations, distributed systems, and ML workflows.

This role involves hands-on Linux systems engineering, GPU driver bring-up, and maintaining the ML software stack (CUDA/ROCm, PyTorch, JAX, vLLM).

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior GPU Cluster Engineer for AI Infrastructure
Senior GPU Cluster Engineer for AI Infrastructure

Sciforium • San Francisco (CA)

On-site
USD 150,000 - 220,000
Medical, dental, and vision insurance
401k plan
Daily lunch, snacks, and beverages
+2
Senior HPC & GPU Infrastructure Engineer
Senior HPC & GPU Infrastructure Engineer

Sciforium • San Francisco (CA)

On-site
USD 150,000 - 220,000
Medical, dental, and vision
401k plan
Lunch and beverages
+2
Senior GPU Compute Cluster Engineer
Senior GPU Compute Cluster Engineer

Inferact Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 200,000 - 400,000
Health benefits
Dental benefits
Vision benefits
+1
Senior GPU Infrastructure Engineer — HPC & Clusters
Senior GPU Infrastructure Engineer — HPC & Clusters

Prime Intellect AI • San Francisco (CA)

On-site
USD 150,000 - 300,000
GPU Cluster Architect: Scalable AI Platform
GPU Cluster Architect: Scalable AI Platform

Sciforium • San Francisco (CA)

On-site
USD 190,000 - 270,000
Medical insurance
401k plan
Daily meals/snacks
+2
GPU Cluster Engineer, Systems & Platform
GPU Cluster Engineer, Systems & Platform

Sciforium • San Francisco (CA)

On-site
USD 150,000 - 220,000
Medical, dental, and vision insurance
401k plan
Daily lunch, snacks, and beverages
+2
Senior GPU Cluster Networking Engineer (RDMA/InfiniBand)
Senior GPU Cluster Networking Engineer (RDMA/InfiniBand)

Sciforium • San Francisco (CA)

On-site
USD 170,000 - 230,000
Senior HPC & GPU Cluster Architect — Scale & Automate
Senior HPC & GPU Cluster Architect — Scale & Automate

San Francisco Compute Company • San Francisco (CA)

Hybrid
USD 120,000 - 160,000
Generous equity grant
Competitive salary
Visa sponsorship
+6
Senior HPC & GPU Cluster Architect
Senior HPC & GPU Cluster Architect

The Consensus • San Francisco (CA)

Hybrid
USD 180,000 - 240,000
Visa sponsorships
401(k) retirement matching
Medical, dental & vision insurance
+2
Senior GPU Infra Lead: Slurm, Kubernetes & Platform
Senior GPU Infra Lead: Slurm, Kubernetes & Platform

Jobgether SRL • United States

Remote
USD 170,000 - 250,000