HPC Infrastructure Engineer

Arcadia

San Francisco (CA)

On-site

USD 180,000 - 260,000

Full time

2 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Arcadia is hiring an HPC Infrastructure Engineer to design, deploy, and operate the foundation of its GPU and TPU cloud. You will work across bare-metal servers, high-speed networking, storage, and automation to deliver a reliable compute platform for customers.

This is a hands-on role for someone who loves cross-domain work and building infrastructure from the ground up in a seed-stage neocloud environment.

Qualifications

  • Experience building or operating production HPC, supercomputing, or large-scale bare-metal infrastructure.
  • Strong Linux systems-engineering and debugging skills.
  • Experience automating physical server provisioning and configuration.
  • Experience with Slurm, Kubernetes, or another distributed workload scheduler.
  • Understanding of high-performance networking, distributed storage, and accelerator-based computing.
  • Experience with monitoring, incident response, performance analysis, and production reliability.
  • Proficiency in Python, Go, Bash, or another infrastructure automation language.
  • Ability to diagnose problems across layers rather than treating compute, networking, storage, and software as separate systems.
  • Comfort working directly with hardware, vendors, data-center teams, and customers in an early-stage environment.

Responsibilities

  • Design, deploy, and operate production TPU clusters
  • Provision and manage Linux-based bare-metal servers at scale
  • Build automated workflows for server installation, configuration, upgrades, and recovery
  • Deploy and operate cluster schedulers such as Slurm or Kubernetes
  • Integrate high-performance networking using InfiniBand or RoCE/RDMA
  • Build and operate high‑throughput storage for distributed AI workloads
  • Monitor cluster health, accelerator utilization, network performance, storage performance, and job reliability
  • Diagnose failures across GPUs, servers, firmware, networks, storage, schedulers, and customer workloads
  • Improve cluster utilization, training performance, fault tolerance, and recovery time
  • Establish production practices for change management, incident response, capacity management, and operational readiness
  • Partner with data-center, network, platform, security, and customer-facing teams to launch new capacity
  • Evaluate and manage infrastructure vendors, hardware suppliers, and technical partners
  • Participate in an on-call rotation as the production platform grows

Skills

Linux systems engineering
HPC infrastructure
Slurm
Kubernetes
Python
Go
Bash
Automation
Networking (InfiniBand)
Diagnostics & debugging

Tools

Ansible
Terraform
Prometheus
Grafana
Lustre
Ceph
GPFS
IPMI

Job description

A seed-stage neocloud company is hiring an HPC Infrastructure Engineer to build and operate the foundation of its GPU and TPU cloud. You'll work across bare-metal systems, high-speed networking, cluster scheduling, storage, automation, and reliability, turning a large fleet of accelerators into a high-performance, dependable computing platform for customers. This is a hands‑on role for someone who loves working across hardware and software, diagnosing tough performance problems, and building infrastructure from the ground up.

What You'll Do
  • Design, deploy, and operate production TPU clusters
  • Provision and manage Linux-based bare-metal servers at scale
  • Build automated workflows for server installation, configuration, upgrades, and recovery
  • Deploy and operate cluster schedulers such as Slurm or Kubernetes
  • Integrate high-performance networking using InfiniBand or RoCE/RDMA
  • Build and operate high‑throughput storage for distributed AI workloads
  • Monitor cluster health, accelerator utilization, network performance, storage performance, and job reliability
  • Diagnose failures across GPUs, servers, firmware, networks, storage, schedulers, and customer workloads
  • Improve cluster utilization, training performance, fault tolerance, and recovery time
  • Establish production practices for change management, incident response, capacity management, and operational readiness
  • Partner with data-center, network, platform, security, and customer-facing teams to launch new capacity
  • Evaluate and manage infrastructure vendors, hardware suppliers, and technical partners
  • Participate in an on-call rotation as the production platform grows
What We're Looking For
  • Experience building or operating production HPC, supercomputing, or large-scale bare-metal infrastructure
  • Strong Linux systems-engineering and debugging skills
  • Experience automating physical server provisioning and configuration
  • Experience with Slurm, Kubernetes, or another distributed workload scheduler
  • Understanding of high-performance networking, distributed storage, and accelerator-based computing
  • Experience with monitoring, incident response, performance analysis, and production reliability
  • Proficiency in Python, Go, Bash, or another infrastructure automation language
  • Ability to diagnose problems across layers rather than treating compute, networking, storage, and software as separate systems
  • Comfort working directly with hardware, vendors, data-center teams, and customers in an early-stage environment
Especially Valuable
  • NVIDIA GPU clusters, TPU systems, CUDA, NCCL, or collective communications
  • InfiniBand, RoCEv2, RDMA, GPUDirect, or high-performance Ethernet
  • Slurm administration and scheduling optimization
  • Kubernetes for accelerator workloads
  • PXE, Redfish, IPMI, BMCs, and bare-metal lifecycle management
  • Ansible, Terraform, or other infrastructure-as-code systems
  • Lustre, Spectrum Scale/GPFS, Ceph, BeeGFS, or other distributed storage systems
  • Prometheus, Grafana, OpenTelemetry, or infrastructure observability tooling
  • GPU health monitoring, firmware management, and hardware failure diagnosis
  • Distributed AI training and inference workloads
  • Performance benchmarking and optimization across compute, network, and storage
  • Experience at a neocloud, hyperscaler, AI lab, national laboratory, or HPC center
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Member of Technical Staff - AI Infrastructure
Member of Technical Staff - AI Infrastructure

Veeda Innovation • California (MO)

Hybrid
USD 180,000 - 240,000
Infrastructure Engineer
Infrastructure Engineer

Acceler8 Talent • San Francisco (CA)

On-site
USD 150,000 - 390,000
Cluster Engineer
Cluster Engineer

STN Inc • San Francisco (CA)

On-site
USD 180,000 - 240,000
GPU / HPC Consultant
GPU / HPC Consultant

Arke • United States

On-site
USD 120,000 - 180,000
Head of AI Data Center Infrastructure Platforms and Software
Head of AI Data Center Infrastructure Platforms and Software

Summit Group Solutions, LLC • United States

On-site
USD 150,000 - 350,000
SRE / Platform Engineer, GPU Infrastructure
SRE / Platform Engineer, GPU Infrastructure

Bake AI • Hillsboro (OR)

On-site
USD 140,000 - 210,000
Sr HPC Hardware Engineer
Sr HPC Hardware Engineer

Career Techniques • Dallas (TX)

Hybrid
USD 120,000 - 180,000
Senior/Staff Software Engineer, Kubernetes Infrastructure
Senior/Staff Software Engineer, Kubernetes Infrastructure

Kindredventures • United States

On-site
USD 140,000 - 190,000
Staff HPC Systems Architect
Staff HPC Systems Architect

Jobtailor • California (MO)

On-site
USD 180,000 - 260,000
Cluster Design
Cluster Design

Blue Signal Search • San Francisco (CA)

On-site
USD 150,000 - 230,000