AI Infra Engineer: GPU Clusters & HPC

Veeda AI

California

On-site

USD 150,000 - 210,000

Full time

36 hours ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Veeda AI in the United States (California) is assembling a small, high‑impact team to architect and operate scalable HPC infrastructure for cutting‑edge multimodal AI workloads. You’ll own GPU clusters, networking, storage, and security, ensuring reliable data delivery and observability across systems.

Ideal candidates have deep Linux and HPC experience, plus automation skills and familiarity with PyTorch/NCCL, Slurm, and cloud PoCs.

Qualifications

  • Bachelor's degree in Computer Science, Computer Engineering, or equivalent hands-on experience in HPC or infrastructure engineering.
  • Strong troubleshooting skills in low-level Linux networking, kernel and PCIe/NUMA tuning, hardware diagnostics, and distributed storage (NFS, NVMe-oF, Lustre, Ceph).
  • Proficient in automation and infrastructure-as-code tools (e.g., Ansible, Terraform, Helm, Python/Bash scripting), and treats cluster configuration as reviewed, version-controlled code.
  • Experience supporting distributed deep learning workloads (PyTorch/NCCL, Ray, DeepSpeed) closely enough to distinguish infrastructure faults from model bugs with controlled benchmarks.
  • Experience administering Linux-based HPC clusters (Slurm/Kubernetes) or cloud PoCs and security policy management.

Responsibilities

  • GPU Cluster Operations: Design, deploy, and operate Slurm-on-Kubernetes GPU clusters, including compute health, interconnects, storage, and scheduling.
  • Access Control and Security: Design and operate identity provisioning and access control for source code, cluster, data storages, and dev tools.
  • Network and Data: Design and maintain distributed data transmission and caching services for just-in-time data delivery to clusters.
  • Observability & Hardware Health: Build telemetry pipelines (Prometheus, Grafana, DCGM, Redfish) to detect hardware issues and coordinate with cloud vendors.
  • Resource Forecasting: Work with Model and Data teams to determine resource needs and lead PoC rounds with vendors.
  • Infra-OPEX: Prepare operational expenses forecasts and provide expert opinion on infra spending decisions.
  • End-to-End Ownership: Lead projects through the full software lifecycle including CI/CD and production observability.

Skills

HPC administration
Linux networking
Automation tooling
Distributed DL workloads
Security policies

Education

Bachelor's in CS/CE or equivalent

Tools

Slurm
Kubernetes
GitLab CI
Terraform
Ansible

Job description

Veeda AI in the United States (California) is assembling a small, high‑impact team to architect and operate scalable HPC infrastructure for cutting‑edge multimodal AI workloads. You’ll own GPU clusters, networking, storage, and security, ensuring reliable data delivery and observability across systems.

Ideal candidates have deep Linux and HPC experience, plus automation skills and familiarity with PyTorch/NCCL, Slurm, and cloud PoCs.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI Infra Engineer – HPC Clusters & Slurm/K8s
AI Infra Engineer – HPC Clusters & Slurm/K8s

Veeda AI • Seattle (WA)

On-site
USD 180,000 - 240,000
Senior AI Infrastructure Engineer - GPU HPC & Slurm Ops
Senior AI Infrastructure Engineer - GPU HPC & Slurm Ops

Veeda Innovation • California (MO)

Hybrid
USD 180,000 - 240,000
Senior AI GPU Infra Architect (Remote, Hyperscale HPC)
Senior AI GPU Infra Architect (Remote, Hyperscale HPC)

5C Group • United States

Remote
USD 120,000 - 150,000
Member of Technical Staff - AI Infrastructure
Member of Technical Staff - AI Infrastructure

Veeda AI • California

On-site
USD 150,000 - 210,000
ML Performance Engineer – Multi-Node Training & Kernels
ML Performance Engineer – Multi-Node Training & Kernels

Veeda AI • California

On-site
USD 180,000 - 240,000
Senior HPC Architect: Large-Scale GPU AI Infra (Equity)
Senior HPC Architect: Large-Scale GPU AI Infra (Equity)

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Equity
Comprehensive benefits
Member of Technical Staff - AI Infrastructure
Member of Technical Staff - AI Infrastructure

Veeda AI • Seattle (WA)

On-site
USD 180,000 - 240,000
ML Performance Engineer: Distributed Training & Inference
ML Performance Engineer: Distributed Training & Inference

Veeda AI • Seattle (WA)

On-site
USD 180,000 - 260,000
Senior AI Infra Engineer - GPU Cluster Architect for Scale
Senior AI Infra Engineer - GPU Cluster Architect for Scale

Telemetry Today LLC • Costa Mesa (CA)

On-site
USD 166,000 - 220,000
Senior AI Infrastructure Lead: GPU Clusters & LLMs
Senior AI Infrastructure Lead: GPU Clusters & LLMs

Cadence Design Systems • San Jose (CA)

On-site
USD 137,000 - 254,000