AI Infrastructure Engineer: HPC GPU Clusters

Veeda AI

Seattle (WA)

On-site

USD 180,000 - 240,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Veeda AI is building the next generation of multimodal foundation world models for Physical AI. The Member of Technical Staff - AI Infrastructure will design, deploy, and operate high-performance GPU clusters running Slurm and Kubernetes.

You will own interconnect fabric, storage pipelines, and observability, ensuring reliability for large-scale distributed DL workloads and real-time data streaming.

Qualifications

  • Bachelor's degree in CS/CE or equivalent HPC/infrastructure experience.
  • Hands-on Linux HPC cluster administration (Slurm or Kubernetes).
  • Strong low-level Linux networking, kernel/PCIe/NUMA tuning, hardware diagnostics.
  • Proficiency with automation and IaC tools; code-driven cluster configuration.
  • Experience with distributed DL workloads (PyTorch/NCCL, Ray, DeepSpeed).

Responsibilities

  • Design, deploy, and operate bare-metal GPU clusters with Slurm and Kubernetes.
  • Tune Slurm topology, fair-share, QoS, and preemption for long runs.
  • Own the interconnect fabric (InfiniBand, RoCEv2) and validate with nccl-tests.
  • Run high-throughput storage (Lustre, WEKA, Ceph) with NVMe tiers.
  • Build telemetry pipelines (Prometheus, Grafana, DCGM) to monitor hardware health.

Skills

HPC cluster administration
Slurm / Kubernetes
Networking / kernel tuning
Automation / IaC
Distributed DL workloads

Education

Bachelor's degree in CS/CE or equivalent HPC experience

Tools

Ansible
Terraform
Helm
Python
NFS/Lustre/Ceph

Job description

Veeda AI is building the next generation of multimodal foundation world models for Physical AI. The Member of Technical Staff - AI Infrastructure will design, deploy, and operate high-performance GPU clusters running Slurm and Kubernetes.

You will own interconnect fabric, storage pipelines, and observability, ensuring reliability for large-scale distributed DL workloads and real-time data streaming.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Infrastructure Architect for HPC & GPU Clusters
AI Infrastructure Architect for HPC & GPU Clusters

Veeda • California (MO)

On-site
USD 150,000 - 200,000
Member of Technical Staff - AI Infrastructure
Member of Technical Staff - AI Infrastructure

Veeda AI • Seattle (WA)

On-site
USD 180,000 - 240,000
Member of Technical Staff - AI Infrastructure
Member of Technical Staff - AI Infrastructure

Veeda • California (MO)

On-site
USD 150,000 - 200,000
Staff Data Engineer, Multimodal AI Pipelines
Staff Data Engineer, Multimodal AI Pipelines

Veeda • California (MO)

On-site
USD 130,000 - 180,000
Senior AI Infra Engineer: Large-Scale GPU & HPC
Senior AI Infra Engineer: Large-Scale GPU & HPC

Anduril Industries • Costa Mesa (CA)

On-site
USD 166,000 - 220,000
Equity grants
Benefits package
Senior AI GPU Cluster Architect
Senior AI GPU Cluster Architect

STN Inc • San Francisco (CA)

On-site
USD 180,000 - 240,000
AI Infra Engineer: GPU Clusters, Automation & HPC
AI Infra Engineer: GPU Clusters, Automation & HPC

Accenture • Detroit (MI)

On-site
USD 110,000 - 210,000
Medical benefits
Dental benefits
401(k) plan
AI Infrastructure Architect — Scalable GPU Compute
AI Infrastructure Architect — Scalable GPU Compute

EngineersOfAI • Sunnyvale (CA)

On-site
USD 150,000 - 200,000
AI Infrastructure Engineer - GPU & Kubernetes
AI Infrastructure Engineer - GPU & Kubernetes

HCLTech • California (MO)

On-site
USD 150,000 - 210,000
Medical Insurance
Dental Insurance
Vision Insurance
+2
Senior AI Training Infra Engineer — GPU Clusters
Senior AI Training Infra Engineer — GPU Clusters

Designworks Talent • Bellevue (WA)

Hybrid
USD 180,000 - 300,000
Medical, dental, vision insurance
401(k) with company match
Paid holidays