Senior GPU Cluster Engineer for AI Infrastructure

Sciforium

San Francisco (CA)

On-site

USD 150,000 - 220,000

Full time

13 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Medical, dental, and vision insurance
401k plan
Daily lunch, snacks, and beverages
Flexible time off
Competitive salary and equity

Job summary

Sciforium is seeking a GPU Cluster Engineer to own the software stack for high-performance GPU clusters. You will define production-ready node images, validate hardware with automated tests, and manage fleet upgrades while ensuring consistent, fast, and scalable performance.

You will work with two stakeholder groups—foundation model teams and model serving teams—deploying in a Kubernetes/Slurm environment with advanced NVIDIA/ROCm stacks and IaC tooling.

Qualifications

  • Bachelor's or Master's in Computer Science, Computer Engineering, Electrical Engineering, or related field.
  • 5+ years in systems/infrastructure with GPU cluster or ML infra experience.
  • Deep Linux internals expertise: kernel modules, DKMS, systemd, cgroups, NUMA.
  • Experience with NVIDIA CUDA and/or ROCm driver stacks on modern accelerators.
  • Production Kubernetes experience with GPU workloads and HPC schedulers.

Responsibilities

  • Own node software definition from base OS to production-ready GPU nodes.
  • Build automated acceptance suites and validation checks for nodes.
  • Maintain fleet with upgrades and topology-aware scheduling.
  • Automate detection of unhealthy nodes and manage re-imaging workflows.

Skills

Linux internals
Python
Bash
GPU cluster engineering
NVIDIA CUDA
ROCm
Kubernetes
Slurm
Ansible/SaltStack

Education

Bachelor's or Master's in CS/CE/EE

Tools

Docker
NVIDIA Container Toolkit
Packer
MaaS
Foreman
Terraform

Job description

Sciforium is seeking a GPU Cluster Engineer to own the software stack for high-performance GPU clusters. You will define production-ready node images, validate hardware with automated tests, and manage fleet upgrades while ensuring consistent, fast, and scalable performance.

You will work with two stakeholder groups—foundation model teams and model serving teams—deploying in a Kubernetes/Slurm environment with advanced NVIDIA/ROCm stacks and IaC tooling.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

GPU Cluster Architect: Scalable AI Platform
GPU Cluster Architect: Scalable AI Platform

Sciforium • San Francisco (CA)

On-site
USD 190,000 - 270,000
Medical insurance
401k plan
Daily meals/snacks
+2
GPU Cluster Engineer, Systems & Platform
GPU Cluster Engineer, Systems & Platform

Sciforium • San Francisco (CA)

On-site
USD 150,000 - 220,000
Medical, dental, and vision insurance
401k plan
Daily lunch, snacks, and beverages
+2
GPU Cluster Ops Engineer: Linux & Hardware
GPU Cluster Ops Engineer: Linux & Hardware

Sciforium • San Francisco (CA)

On-site
USD 120,000 - 160,000
Medical insurance
Dental insurance
Vision insurance
+4
GPU Cluster Engineer, Systems & Platform
GPU Cluster Engineer, Systems & Platform

Sciforium • San Francisco (CA)

On-site
USD 190,000 - 270,000
Medical insurance
401k plan
Daily meals/snacks
+2
GPU Cluster Engineer, Hardware Operations
GPU Cluster Engineer, Hardware Operations

Sciforium • San Francisco (CA)

On-site
USD 120,000 - 160,000
Medical insurance
Dental insurance
Vision insurance
+4
Senior GPU Cluster Networking Engineer (RDMA/InfiniBand)
Senior GPU Cluster Networking Engineer (RDMA/InfiniBand)

Sciforium • San Francisco (CA)

On-site
USD 170,000 - 230,000
Senior AI Infrastructure Engineer — Scalable GPU Clusters
Senior AI Infrastructure Engineer — Scalable GPU Clusters

NVIDIA AI • Santa Clara (CA)

On-site
USD 152,000 - 288,000
Senior GPU Infra Engineer for Distributed AI
Senior GPU Infra Engineer for Distributed AI

Andromeda Cluster • San Francisco (CA)

Hybrid
USD 180,000 - 240,000
Senior Full-Stack Engineer, AI Infra for GPU Clusters
Senior Full-Stack Engineer, AI Infra for GPU Clusters

NVIDIA • California (MO)

On-site
USD 184,000 - 357,000
Equity
Benefits
GPU Cluster Engineer - Scalable AI Infrastructure
GPU Cluster Engineer - Scalable AI Infrastructure

NEURA Robotics • Germany (OH)

On-site
USD 140,000 - 195,000