GPU Cluster Architect: Scalable AI Platform

Sciforium

San Francisco (CA)

On-site

USD 190,000 - 270,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Medical insurance
401k plan
Daily meals/snacks
Flexible time off
Competitive salary and equity

Job summary

Sciforium is seeking a GPU Cluster Engineer to own the software stack of GPU clusters—from kernel tuning to ML framework integration. You will define production-ready nodes, automate bring-up, and maintain fleet health while supporting foundation model training and model serving teams.

You will automate image creation, validation suites, and fleet upgrades, leveraging Kubernetes, Slurm, and container tooling to ensure high performance and reliability across the cluster.

Qualifications

  • 5+ years in systems/infrastructure engineering with GPU cluster/ML infra experience.
  • Bachelor's or Master's in CS/EE or related field.
  • Deep Linux internals: kernel modules, DKMS, NUMA, cgroups, system performance tuning.
  • Hands-on experience with NVIDIA CUDA and/or ROCm driver stacks on modern accelerators.
  • Production Kubernetes experience with GPU workloads; knowledge of HPC schedulers (Slurm/Run:AI).
  • Configuration management with Ansible or SaltStack; Git-based workflows.
  • Image provisioning/tooling (Packer, MaaS, Foreman, Terraform).
  • Experience with distributed filesystems (Lustre, GPFS) and checkpoint I/O optimization.
  • Python and Bash for automation; NCCL/RCCL and PyTorch/JAX runtime familiarity.

Responsibilities

  • Own node software definition from OS bring-up to production-ready GPU nodes.
  • Build automated acceptance suites and validation for fleet readiness.
  • Maintain fleet through kernel/driver/toolkit upgrades with minimal disruption.
  • Automate health checks, drift detection, and repair handoffs.
  • Manage node/configuration via IaC (Ansible/SaltStack) with CI validation and canaries.
  • Develop tooling for cluster operations, health reporting, and workflows.
  • Operate GPU-enabled Kubernetes and Slurm for serving and training workloads.
  • Maintain CUDA/ROCm stacks, container tooling, and reproducible images.

Skills

GPU cluster
Linux internals
Python scripting
Kubernetes
NVIDIA CUDA
Ansible
System performance

Education

Bachelor's or Master's in CS/EE

Tools

NVIDIA Driver Toolkit
ROCm
Packer
Terraform

Job description

Sciforium is seeking a GPU Cluster Engineer to own the software stack of GPU clusters—from kernel tuning to ML framework integration. You will define production-ready nodes, automate bring-up, and maintain fleet health while supporting foundation model training and model serving teams.

You will automate image creation, validation suites, and fleet upgrades, leveraging Kubernetes, Slurm, and container tooling to ensure high performance and reliability across the cluster.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior GPU Cluster Engineer for AI Infrastructure
Senior GPU Cluster Engineer for AI Infrastructure

Sciforium • San Francisco (CA)

On-site
USD 150,000 - 220,000
Medical, dental, and vision insurance
401k plan
Daily lunch, snacks, and beverages
+2
GPU Cluster Engineer, Systems & Platform
GPU Cluster Engineer, Systems & Platform

Sciforium • San Francisco (CA)

On-site
USD 150,000 - 220,000
Medical, dental, and vision insurance
401k plan
Daily lunch, snacks, and beverages
+2
GPU Cluster Engineer, Systems & Platform
GPU Cluster Engineer, Systems & Platform

Sciforium • San Francisco (CA)

On-site
USD 190,000 - 270,000
Medical insurance
401k plan
Daily meals/snacks
+2
Lead GPU HPC Infrastructure Engineer
Lead GPU HPC Infrastructure Engineer

Sciforium • San Francisco (CA)

On-site
USD 180,000 - 240,000
Medical, dental, and vision insurance
401k plan
Daily lunch, snacks, and beverages
+2
Senior GPU Cluster Networking Engineer (RDMA/InfiniBand)
Senior GPU Cluster Networking Engineer (RDMA/InfiniBand)

Sciforium • San Francisco (CA)

On-site
USD 170,000 - 230,000
GPU Cluster Infra Lead — Tech Strategy & Team Growth
GPU Cluster Infra Lead — Tech Strategy & Team Growth

Far Ai • United States

Remote
USD 180,000 - 250,000
Senior GPU Cluster Infra Engineer | Remote
Senior GPU Cluster Infra Engineer | Remote

AISafety • Berkeley (CA)

Hybrid
USD 120,000 - 180,000
Health Insurance
401(k) match
PTO 25 days per year
+3
Senior GPU Infrastructure Engineer — HPC & Clusters
Senior GPU Infrastructure Engineer — HPC & Clusters

Prime Intellect AI • San Francisco (CA)

On-site
USD 150,000 - 300,000
Senior HPC & GPU Infrastructure Engineer
Senior HPC & GPU Infrastructure Engineer

Sciforium • San Francisco (CA)

On-site
USD 180,000 - 240,000
Medical, dental, and vision insurance
401k plan
Daily lunch, snacks, and beverages
+2
GPU Cluster Infra Tech Lead & Team Builder
GPU Cluster Infra Tech Lead & Team Builder

AISafety • Berkeley (CA)

Hybrid
USD 180,000 - 240,000
Health Insurance
401(k) plan
PTO - 25 days per year
+4