AI Infrastructure Architect for HPC & GPU Clusters

Veeda

California (MO)

On-site

USD 150,000 - 200,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Veeda AI is seeking a Member of Technical Staff to architect and operate AI infrastructure for large-scale multimodal models. You will manage bare-metal GPU clusters, scheduling, fabric interconnects, and high-throughput storage to keep training sessions running at scale.

Ideal candidates have deep HPC experience, proficiency with Slurm and Kubernetes, and a track record of building observability pipelines and automation. Join a fast-moving team tackling frontier AI research with direct impact.

Qualifications

  • Bachelor's degree in Computer Science, Computer Engineering, or equivalent hands-on HPC experience.
  • Deep hands-on expertise administering Linux-based HPC clusters with Slurm or Kubernetes.
  • Strong troubleshooting across Linux networking, kernel, PCIe/NUMA tuning, hardware diagnostics.
  • Proficient in automation and IaC tools, with code-reviewed, version-controlled configurations.
  • Experience supporting distributed deep learning workloads (PyTorch/NCCL, Ray, DeepSpeed).

Responsibilities

  • GPU Cluster Operations: Design, deploy, and operate bare-metal GPU clusters where Slurm and Kubernetes share nodes (Slinky, KubeRay, MPI Operator).
  • Scheduler & Topology: Tune Slurm block and topology plugins so a job lands inside one NVLink domain, and set fair-share, QoS, and preemption policy so long training runs and bursty simulation rollouts coexist.
  • Fabric Engineering: Own the interconnect (InfiniBand subnet manager, adaptive routing and SHARP, or RoCEv2 with PFC and ECN tuning), and validate it with nccl-tests before a hang gets blamed on the model.
  • Storage & Data Path: Run high-throughput storage (Lustre, WEKA, Ceph) with NVMe scratch tiers and caching so video datasets stream at line rate and checkpoint writes never stall a run.
  • Observability & Hardware Health: Build the telemetry pipeline (Prometheus, Grafana, DCGM, BMC/Redfish) that catches Xid and ECC errors, thermal throttling, and link flaps, and automate the drain-and-replace that keeps them off live jobs.

Skills

HPC cluster admin
Slurm administration
Kubernetes administration
Linux troubleshooting
GPU hardware expertise
NVIDIA GPUs
PyTorch NCCL
Distributed storage

Education

Bachelor's degree in CS/CE or equivalent

Tools

Ansible
Terraform
Helm
Python scripting
Bash scripting

Job description

Veeda AI is seeking a Member of Technical Staff to architect and operate AI infrastructure for large-scale multimodal models. You will manage bare-metal GPU clusters, scheduling, fabric interconnects, and high-throughput storage to keep training sessions running at scale.

Ideal candidates have deep HPC experience, proficiency with Slurm and Kubernetes, and a track record of building observability pipelines and automation. Join a fast-moving team tackling frontier AI research with direct impact.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Infrastructure Engineer: HPC GPU Clusters
AI Infrastructure Engineer: HPC GPU Clusters

Veeda AI • Seattle (WA)

On-site
USD 180,000 - 240,000
Member of Technical Staff - AI Infrastructure
Member of Technical Staff - AI Infrastructure

Veeda AI • Seattle (WA)

On-site
USD 180,000 - 240,000
Member of Technical Staff - AI Infrastructure
Member of Technical Staff - AI Infrastructure

Veeda • California (MO)

On-site
USD 150,000 - 200,000
Senior AI Infra Engineer: Large-Scale GPU & HPC
Senior AI Infra Engineer: Large-Scale GPU & HPC

Anduril Industries • Costa Mesa (CA)

On-site
USD 166,000 - 220,000
Equity grants
Benefits package
Staff Data Engineer, Multimodal AI Pipelines
Staff Data Engineer, Multimodal AI Pipelines

Veeda • California (MO)

On-site
USD 130,000 - 180,000
Senior AI Scheduling & Orchestration Engineer
Senior AI Scheduling & Orchestration Engineer

Bitdeer Technologies Group • Austin (TX)

On-site
USD 180,000 - 280,000
Senior AI Infrastructure Solutions Architect
Senior AI Infrastructure Solutions Architect

VAST Data • New York (NY)

On-site
USD 150,000 - 200,000
Senior AI GPU Cluster Architect
Senior AI GPU Cluster Architect

STN Inc • San Francisco (CA)

On-site
USD 180,000 - 240,000
Senior AI Infrastructure Engineer — Scale GPU Clusters Remote
Senior AI Infrastructure Engineer — Scale GPU Clusters Remote

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 280,000 - 420,000
Equity options
Health, vision, dental benefits
Unlimited PTO
+2
AI Infrastructure Architect — Scalable GPU Compute
AI Infrastructure Architect — Scalable GPU Compute

EngineersOfAI • Sunnyvale (CA)

On-site
USD 150,000 - 200,000