Senior AI Infrastructure Engineer - GPU HPC & Slurm Ops

Veeda Innovation

California (MO)

Hybrid

USD 180,000 - 240,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Veeda AI is hiring a Member of Technical Staff - AI Infrastructure to design, deploy, and operate GPU clusters (Slurm-on-Kubernetes) with focus on health, interconnect, storage, and scheduling. You will build data pipelines, security controls, and observability to support large-scale AI workloads.

You will collaborate with Model/Data teams, forecast resources, run PoCs with vendors, and own infra lifecycle from specs to production, including on-call support and cost optimization.

Qualifications

  • Bachelor's degree in Computer Science, Computer Engineering, or equivalent hands-on experience in HPC or infrastructure engineering
  • Strong troubleshooting skills below the framework layer: low-level Linux networking, kernel and PCIe/NUMA tuning, hardware diagnostics, and network or distributed storage (NFS, NVMe-oF, Lustre, Ceph)
  • Proficient in automation and infrastructure-as-code tools (e.g., Ansible, Terraform, Helm, Python/Bash scripting) and treats cluster configuration as reviewed, version-controlled code
  • Experience supporting distributed deep learning workloads (PyTorch/NCCL, Ray, DeepSpeed) closely enough to tell an infrastructure fault from a model bug, and able to prove which with a controlled benchmark
  • One of the following: Deep hands-on expertise administering Linux-based HPC clusters running Slurm or Kubernetes (job scheduling, cgroups, fair-share policies, topology configuration, GPU device plugins); Experience working with cloud providers on workload profiling, PoC, resource cost strategies; Experience managing security policies. Hands-on-experience with building security layers from overall strategy to IaC.

Responsibilities

  • GPU Cluster Operations: Design, deploy, and operate Slurm-on-Kubernetes GPU clusters and manage compute, interconnect, storage, and scheduling resources
  • Access Control and Security: Design, deploy, and operate identity provisioning and access control for source code, cluster, knowledge bases, data storages, and dev tools
  • Network and Data: Design and maintain distributed data transmission and caching services to deliver data and checkpoints to clusters efficiently
  • Observability & Hardware Health: Build telemetry pipeline (Prometheus, Grafana, DCGM, BMC/Redfish) to catch errors and work with vendors to resolve issues
  • Resource Forecasting: Determine resource requirements and cadence, lead PoC rounds with vendors for compute/data resources
  • Infra-OPEX: Prepare operational expenses forecast and reports, advise on infra spending decisions
  • End-to-End Ownership: Lead projects through full software lifecycle including specs, implementation, CI/CD, on-call support, observability

Skills

Linux troubleshooting
HPC infra experience
Automation & IaC
DL workloads support

Education

Bachelor's degree in CS/CE or equivalent

Tools

Ansible
Terraform
Helm
Python/Bash scripting

Job description

Veeda AI is hiring a Member of Technical Staff - AI Infrastructure to design, deploy, and operate GPU clusters (Slurm-on-Kubernetes) with focus on health, interconnect, storage, and scheduling. You will build data pipelines, security controls, and observability to support large-scale AI workloads.

You will collaborate with Model/Data teams, forecast resources, run PoCs with vendors, and own infra lifecycle from specs to production, including on-call support and cost optimization.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI Infra Engineer – HPC Clusters & Slurm/K8s
AI Infra Engineer – HPC Clusters & Slurm/K8s

Veeda AI • Seattle (WA)

On-site
USD 180,000 - 240,000
AI Infra Engineer: GPU Clusters & HPC
AI Infra Engineer: GPU Clusters & HPC

Veeda AI • California

On-site
USD 150,000 - 210,000
Member of Technical Staff - AI Infrastructure
Member of Technical Staff - AI Infrastructure

Veeda AI • Seattle (WA)

On-site
USD 180,000 - 240,000
Member of Technical Staff - AI Infrastructure
Member of Technical Staff - AI Infrastructure

Veeda AI • California

On-site
USD 150,000 - 210,000
Senior Slurm & HPC Cluster Engineer (GPU/AI Infra)
Senior Slurm & HPC Cluster Engineer (GPU/AI Infra)

Bitdeer (NASDAQ: BTDR) • Austin (TX)

On-site
USD 150,000 - 190,000
Senior AI GPU Infra Architect (Remote, Hyperscale HPC)
Senior AI GPU Infra Architect (Remote, Hyperscale HPC)

5C Group • United States

Remote
USD 120,000 - 150,000
Member of Technical Staff - AI Infrastructure
Member of Technical Staff - AI Infrastructure

Veeda Innovation • California (MO)

Hybrid
USD 180,000 - 240,000
Senior AI Datacenter Software Engineer (Kubernetes/Slurm)
Senior AI Datacenter Software Engineer (Kubernetes/Slurm)

BranchFactor • Austin (TX), Northern (KY)

Hybrid
USD 184,000 - 357,000
Equity
Benefits
Senior GPU Infrastructure Engineer — HPC & Clusters
Senior GPU Infrastructure Engineer — HPC & Clusters

Prime Intellect AI • San Francisco (CA)

On-site
USD 150,000 - 300,000
Senior AI Infrastructure Solutions Architect
Senior AI Infrastructure Solutions Architect

VAST Data • New York (NY)

On-site
USD 150,000 - 200,000