GPU Cluster Architect: Scalable AI Platform

Sciforium

San Francisco (CA)

On-site

USD 190,000 - 270,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Medical insurance
401k plan
Daily meals/snacks
Flexible time off
Competitive salary and equity

Job summary

Sciforium is seeking a GPU Cluster Engineer to own the software stack of GPU clusters—from kernel tuning to ML framework integration. You will define production-ready nodes, automate bring-up, and maintain fleet health while supporting foundation model training and model serving teams.

You will automate image creation, validation suites, and fleet upgrades, leveraging Kubernetes, Slurm, and container tooling to ensure high performance and reliability across the cluster.

Qualifications

  • 5+ years in systems/infrastructure engineering with GPU cluster/ML infra experience.
  • Bachelor's or Master's in CS/EE or related field.
  • Deep Linux internals: kernel modules, DKMS, NUMA, cgroups, system performance tuning.
  • Hands-on experience with NVIDIA CUDA and/or ROCm driver stacks on modern accelerators.
  • Production Kubernetes experience with GPU workloads; knowledge of HPC schedulers (Slurm/Run:AI).
  • Configuration management with Ansible or SaltStack; Git-based workflows.
  • Image provisioning/tooling (Packer, MaaS, Foreman, Terraform).
  • Experience with distributed filesystems (Lustre, GPFS) and checkpoint I/O optimization.
  • Python and Bash for automation; NCCL/RCCL and PyTorch/JAX runtime familiarity.

Responsibilities

  • Own node software definition from OS bring-up to production-ready GPU nodes.
  • Build automated acceptance suites and validation for fleet readiness.
  • Maintain fleet through kernel/driver/toolkit upgrades with minimal disruption.
  • Automate health checks, drift detection, and repair handoffs.
  • Manage node/configuration via IaC (Ansible/SaltStack) with CI validation and canaries.
  • Develop tooling for cluster operations, health reporting, and workflows.
  • Operate GPU-enabled Kubernetes and Slurm for serving and training workloads.
  • Maintain CUDA/ROCm stacks, container tooling, and reproducible images.

Skills

GPU cluster
Linux internals
Python scripting
Kubernetes
NVIDIA CUDA
Ansible
System performance

Education

Bachelor's or Master's in CS/EE

Tools

NVIDIA Driver Toolkit
ROCm
Packer
Terraform

Job description

Sciforium is seeking a GPU Cluster Engineer to own the software stack of GPU clusters—from kernel tuning to ML framework integration. You will define production-ready nodes, automate bring-up, and maintain fleet health while supporting foundation model training and model serving teams.

You will automate image creation, validation suites, and fleet upgrades, leveraging Kubernetes, Slurm, and container tooling to ensure high performance and reliability across the cluster.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior GPU Cluster Engineer for AI Infrastructure
Senior GPU Cluster Engineer for AI Infrastructure

Sciforium • San Francisco (CA)

On-site
USD 150,000 - 220,000
Medical, dental, and vision insurance
401k plan
Daily lunch, snacks, and beverages
+2
GPU Cluster Ops Engineer: Linux & Hardware
GPU Cluster Ops Engineer: Linux & Hardware

Sciforium • San Francisco (CA)

On-site
USD 120,000 - 160,000
Medical insurance
Dental insurance
Vision insurance
+4
GPU Cluster Engineer, Hardware Operations
GPU Cluster Engineer, Hardware Operations

Sciforium • San Francisco (CA)

On-site
USD 120,000 - 160,000
Medical insurance
Dental insurance
Vision insurance
+4
GPU Cluster Engineer - Scalable AI Infrastructure
GPU Cluster Engineer - Scalable AI Infrastructure

NEURA Robotics • Germany (OH)

On-site
USD 140,000 - 195,000
GPU Cluster Engineer, Systems & Platform
GPU Cluster Engineer, Systems & Platform

Sciforium • San Francisco (CA)

On-site
USD 150,000 - 220,000
Medical, dental, and vision insurance
401k plan
Daily lunch, snacks, and beverages
+2
GPU Cluster Infrastructure Engineer
GPU Cluster Infrastructure Engineer

PVH (Tommy Hilfiger/Calvin Klein) • United States

On-site
USD 140,000 - 210,000
Equity in unicorn-stage company
100% premiums covered for medical, den
401(k) matching up to 4%
+2
GPU Cluster Engineer, Systems & Platform
GPU Cluster Engineer, Systems & Platform

Sciforium • San Francisco (CA)

On-site
USD 190,000 - 270,000
Medical insurance
401k plan
Daily meals/snacks
+2
GPU Cluster Infra Engineer - Reliability & Automation
GPU Cluster Infra Engineer - Reliability & Automation

Doist • San Francisco (CA)

On-site
USD 150,000 - 210,000
Equity
Health benefits
401k matching
+2
Senior GPU Cluster Networking Engineer (RDMA/InfiniBand)
Senior GPU Cluster Networking Engineer (RDMA/InfiniBand)

Sciforium • San Francisco (CA)

On-site
USD 170,000 - 230,000
Senior AI Infrastructure Engineer — Scalable GPU Clusters
Senior AI Infrastructure Engineer — Scalable GPU Clusters

NVIDIA AI • Santa Clara (CA)

On-site
USD 152,000 - 288,000