HPC/ML Infrastructure Engineer

Spellbrush

San Francisco (CA)

On-site

USD 120,000 - 150,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Spellbrush in San Francisco seeks an experienced HPC Infrastructure Engineer to manage and operate one of the largest anime AI training clusters globally. You will collaborate directly with researchers to ensure system performance and manage SLURM jobs.

The role requires familiarity with cutting-edge HPC software, strong Linux skills, and the ability to work in a fast-paced environment. Preference is given to candidates located in the Bay Area for on-site collaboration, with visa sponsorship available.

Qualifications

  • Experience in large-scale GPU systems management.
  • Ability to install and configure SLURM and handle complex HPC environments.
  • Strong Linux sysadmin skills including directory management.

Responsibilities

  • Lead the administration of the largest anime AI training cluster.
  • Ensure SLURM jobs are running and manage network configurations.
  • Assist AI researchers in training anime models.

Skills

HPC software landscape familiarity
Linux sysadmin skills
Experience with SLURM
Working with GPUs
Networking configurations

Tools

Kubernetes
Ansible
Grafana
Prometheus
Ceph

Job description

We’re looking for an experienced HPC infrastructure engineer to lead bringup, administration, and operations on is probably the largest anime AI training cluster in the world. You’ll serve as the bridge between our researchers and the bare GPU machines, helping to make sure that SLURM jobs are running, parallel filesystems are serving, network is transmitting, and that the anime models are training.

You may be a good fit if:
You love anime and the anime aesthetic.

This probably one of the only jobs in the world where you will get to combine your love of anime and large-scale GPU systems.

You’re familiar with the modern HPC software landscape

Once upon a time, our team could install SLURM on a few bare metal nodes and get away with it. Now the landscape has become unbelievable complex, with SLURM deploys through Slinky on K8s, provisioning through warewulf/MAAS/ansible, filesystems through WEKA/VAST/Ceph, VPN and access through tailscale, and monitoring via the Grafana/Prometheus stack. We’re looking for someone with relevant experience up and down the stack (and maybe a papercut or two to show for it!)

As well as the traditional sysadmin landscape

Bringing up and managing cluster still requires good old linux sysadmin skills, including wrangling ldap, triaging dmesg, and setting sticky bits on directories for misbehaving users and tools.

You're not afraid of physical computers

We’re building out edge datacenters and our CEO is still personally racking, stacking, and provisioning HGX-based nodes in our living room. Also his VLAN design sucks and he’s bad at fiber routing. Please send help.

And you're comfortable working on small, fast-paced teams.

We currently have a very tiny research team, and you’ll be directly helping some of the AI researchers in the world train the best anime image model in the world.

We also believe in the unmatched speed of in-person teams, and prefer on-site collaboration in either our primary research office in Tokyo (downtown Akihabara), or San Francisco (dogpatch!). Bay area is strongly preferred as we have physical hardware in the Bay Area. Visa sponsorships are available.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

HPC/ML Infra Engineer — Anime AI Training Cluster
HPC/ML Infra Engineer — Anime AI Training Cluster

Spellbrush • San Francisco (CA)

On-site
USD 120,000 - 150,000
AI Infrastructure Engineer
AI Infrastructure Engineer

Visa Hunt • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Equity
Health insurance
Dental insurance
+1
Member of Technical Staff (AI Infrastructure Engineer)
Member of Technical Staff (AI Infrastructure Engineer)

Pantera Capital • Palo Alto (CA)

Hybrid
USD 190,000 - 250,000
Comprehensive health insurance
Dental and vision insurance
401(k) plan
+1
Member of Technical Staff (AI Infrastructure Engineer)
Member of Technical Staff (AI Infrastructure Engineer)

Perplexity • San Francisco (CA)

On-site
USD 120,000 - 150,000
HPC AI Systems Administrator
HPC AI Systems Administrator

MRE Consulting • Houston (TX)

On-site
USD 95,000 - 140,000
Senior Site Reliability Engineer (SRE) - AI Inftastructure
Senior Site Reliability Engineer (SRE) - AI Inftastructure

Hamilton Barnes Associates Limited • San Francisco (CA)

On-site
USD 270,000 - 330,000
Equity
Software Engineer
Software Engineer

Visa Hunt • San Francisco (CA), Northern (KY)

Hybrid
USD 140,000 - 200,000
Infrastructure, Large-scale Training San Jose
Infrastructure, Large-scale Training San Jose

Hark, Inc. • San Jose (CA)

On-site
USD 180,000 - 450,000
Head of AI Data Center Infrastructure Platforms and Software
Head of AI Data Center Infrastructure Platforms and Software

Summit Group Solutions, LLC • United States

On-site
USD 150,000 - 350,000
HPC/ GPU Hardware Engineer
HPC/ GPU Hardware Engineer

The San Francisco Compute Company • San Francisco (CA)

Hybrid
USD 180,000 - 260,000
Generous equity grant
Retirement matching
Comprehensive medical, dental, and vision insurance
+5