HPC/ML Infra Engineer — Anime AI Training Cluster

Spellbrush

San Francisco (CA)

On-site

USD 120,000 - 150,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Spellbrush in San Francisco seeks an experienced HPC Infrastructure Engineer to manage and operate one of the largest anime AI training clusters globally. You will collaborate directly with researchers to ensure system performance and manage SLURM jobs.

The role requires familiarity with cutting-edge HPC software, strong Linux skills, and the ability to work in a fast-paced environment. Preference is given to candidates located in the Bay Area for on-site collaboration, with visa sponsorship available.

Qualifications

  • Experience in large-scale GPU systems management.
  • Ability to install and configure SLURM and handle complex HPC environments.
  • Strong Linux sysadmin skills including directory management.

Responsibilities

  • Lead the administration of the largest anime AI training cluster.
  • Ensure SLURM jobs are running and manage network configurations.
  • Assist AI researchers in training anime models.

Skills

HPC software landscape familiarity
Linux sysadmin skills
Experience with SLURM
Working with GPUs
Networking configurations

Tools

Kubernetes
Ansible
Grafana
Prometheus
Ceph

Job description

Spellbrush in San Francisco seeks an experienced HPC Infrastructure Engineer to manage and operate one of the largest anime AI training clusters globally. You will collaborate directly with researchers to ensure system performance and manage SLURM jobs.

The role requires familiarity with cutting-edge HPC software, strong Linux skills, and the ability to work in a fast-paced environment. Preference is given to candidates located in the Bay Area for on-site collaboration, with visa sponsorship available.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

HPC/ML Infrastructure Engineer
HPC/ML Infrastructure Engineer

Spellbrush • San Francisco (CA)

On-site
USD 120,000 - 150,000
Senior Data Pipeline & AI Infra Engineer
Senior Data Pipeline & AI Infra Engineer

Visa Hunt • San Francisco (CA), Northern (KY)

Hybrid
USD 140,000 - 200,000
Member of Technical Staff (AI Infrastructure Engineer)
Member of Technical Staff (AI Infrastructure Engineer)

Perplexity • California (MO)

On-site
USD 140,000 - 190,000
Member of Technical Staff (AI Infrastructure Engineer)
Member of Technical Staff (AI Infrastructure Engineer)

Pantera Capital • Palo Alto (CA)

Hybrid
USD 190,000 - 250,000
Comprehensive health insurance
Dental and vision insurance
401(k) plan
+1
AI Infrastructure Engineer
AI Infrastructure Engineer

Visa Hunt • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Equity
Health insurance
Dental insurance
+1
HPC Cluster Engineer — AI/ML & OpenShift Infra
HPC Cluster Engineer — AI/ML & OpenShift Infra

Linuxconfig • Springfield (VA)

Hybrid
USD 140,000 - 185,000
Lead Large-Scale GPU Cluster Engineer for AI Research
Lead Large-Scale GPU Cluster Engineer for AI Research

Linuxcareers • San Francisco (CA)

On-site
USD 120,000 - 180,000
AI Infra Engineer: Scale ML Clusters with Kubernetes
AI Infra Engineer: Scale ML Clusters with Kubernetes

Perplexity • California (MO)

On-site
USD 140,000 - 190,000
HPC Infrastructure & AI Compute Cluster Engineer
HPC Infrastructure & AI Compute Cluster Engineer

INflow • Springfield (VA)

On-site
USD 140,000 - 185,000
Senior AI GPU Cluster Architect
Senior AI GPU Cluster Architect

STN Inc • San Francisco (CA)

On-site
USD 180,000 - 240,000