Senior GPU Cluster Engineer for AI Infrastructure

Sciforium

San Francisco (CA)

On-site

USD 150,000 - 220,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Medical, dental, and vision insurance
401k plan
Daily lunch, snacks, and beverages
Flexible time off
Competitive salary and equity

Job summary

Sciforium is seeking a GPU Cluster Engineer to own the software stack for high-performance GPU clusters. You will define production-ready node images, validate hardware with automated tests, and manage fleet upgrades while ensuring consistent, fast, and scalable performance.

You will work with two stakeholder groups—foundation model teams and model serving teams—deploying in a Kubernetes/Slurm environment with advanced NVIDIA/ROCm stacks and IaC tooling.

Qualifications

  • Bachelor's or Master's in Computer Science, Computer Engineering, Electrical Engineering, or related field.
  • 5+ years in systems/infrastructure with GPU cluster or ML infra experience.
  • Deep Linux internals expertise: kernel modules, DKMS, systemd, cgroups, NUMA.
  • Experience with NVIDIA CUDA and/or ROCm driver stacks on modern accelerators.
  • Production Kubernetes experience with GPU workloads and HPC schedulers.

Responsibilities

  • Own node software definition from base OS to production-ready GPU nodes.
  • Build automated acceptance suites and validation checks for nodes.
  • Maintain fleet with upgrades and topology-aware scheduling.
  • Automate detection of unhealthy nodes and manage re-imaging workflows.

Skills

Linux internals
Python
Bash
GPU cluster engineering
NVIDIA CUDA
ROCm
Kubernetes
Slurm
Ansible/SaltStack

Education

Bachelor's or Master's in CS/CE/EE

Tools

Docker
NVIDIA Container Toolkit
Packer
MaaS
Foreman
Terraform

Job description

Sciforium is seeking a GPU Cluster Engineer to own the software stack for high-performance GPU clusters. You will define production-ready node images, validate hardware with automated tests, and manage fleet upgrades while ensuring consistent, fast, and scalable performance.

You will work with two stakeholder groups—foundation model teams and model serving teams—deploying in a Kubernetes/Slurm environment with advanced NVIDIA/ROCm stacks and IaC tooling.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

GPU Cluster Architect: Scalable AI Platform
GPU Cluster Architect: Scalable AI Platform

Sciforium • San Francisco (CA)

On-site
USD 190,000 - 270,000
Medical insurance
401k plan
Daily meals/snacks
+2
Lead GPU HPC Infrastructure Engineer
Lead GPU HPC Infrastructure Engineer

Sciforium • San Francisco (CA)

On-site
USD 180,000 - 240,000
Medical, dental, and vision insurance
401k plan
Daily lunch, snacks, and beverages
+2
GPU Cluster Engineer, Systems & Platform
GPU Cluster Engineer, Systems & Platform

Sciforium • San Francisco (CA)

On-site
USD 150,000 - 220,000
Medical, dental, and vision insurance
401k plan
Daily lunch, snacks, and beverages
+2
GPU Cluster Engineer, Systems & Platform
GPU Cluster Engineer, Systems & Platform

Sciforium • San Francisco (CA)

On-site
USD 190,000 - 270,000
Medical insurance
401k plan
Daily meals/snacks
+2
Senior GPU Cluster Networking Engineer (RDMA/InfiniBand)
Senior GPU Cluster Networking Engineer (RDMA/InfiniBand)

Sciforium • San Francisco (CA)

On-site
USD 170,000 - 230,000
Senior GPU Cluster Infra Engineer | Remote
Senior GPU Cluster Infra Engineer | Remote

AISafety • Berkeley (CA)

Hybrid
USD 120,000 - 180,000
Health Insurance
401(k) match
PTO 25 days per year
+3
Senior HPC & GPU Infrastructure Engineer
Senior HPC & GPU Infrastructure Engineer

Sciforium • San Francisco (CA)

On-site
USD 180,000 - 240,000
Medical, dental, and vision insurance
401k plan
Daily lunch, snacks, and beverages
+2
Senior GPU Infrastructure Engineer — HPC & Clusters
Senior GPU Infrastructure Engineer — HPC & Clusters

Prime Intellect AI • San Francisco (CA)

On-site
USD 150,000 - 300,000
Senior GPU Cluster Architect for AI Infra at Scale
Senior GPU Cluster Architect for AI Infra at Scale

Partner Company • United States

Remote
USD 184,000 - 318,000
Medical insurance
Dental insurance
Vision insurance
+1
Senior GPU Infra Lead: Slurm, Kubernetes & Platform
Senior GPU Infra Lead: Slurm, Kubernetes & Platform

Jobgether SRL • United States

Remote
USD 170,000 - 250,000