Senior GPU Infra Architect for Scalable AI Compute

AI Chopping Block

Costa Mesa, Northern (CA, KY)

Hybrid

USD 166,000 - 220,000

Full time

5 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Anduril Industries is seeking a Senior AI Infrastructure Engineer to own the stability and scalability of large-scale GPU training. You will drive automation and optimize Kubernetes/Run:AI for research and product teams.

This role requires 10+ years in HPC/datacenter infra, hands-on GPU hardware experience, and eligibility for a U.S. Top Secret clearance. You’ll build self-healing systems and improve interconnects to support massive-scale compute.

Qualifications

  • 10+ years in hands-on infrastructure, HPC, or datacenter engineering supporting GPU compute at scale.
  • Experience with H200/B200/B300 GPUs: bring up, cabling, firmware/driver management.
  • Experience with high-performance interconnects (NVLink, InfiniBand, RoCE, Spectrum-X) in clusters of GPUs.
  • Experience with high-performance parallel storage (VAST, DDN, Weka, Lustre).
  • Kubernetes required; Run:AI or similar GPU scheduling experience preferred.
  • Strong automation background with repeatable deployment pipelines.
  • Able to lift/move 50+ lbs and perform physical datacenter work; U.S. Top Secret clearance eligible.

Responsibilities

  • Rack, stack, cable, and bring up GPU compute including topology, power, cooling, firmware, and burn-in validation.
  • Build and tune interconnect fabric (NVLink, InfiniBand, RoCE, Spectrum-X) for hundreds of GPUs.
  • Integrate high-performance parallel storage to sustain throughput for distributed training and datasets.
  • Automate cluster deployment and configuration end to end, IaC for bring up and driver management.
  • Operate and extend Kubernetes/Run:AI for GPU scheduling and multi-tenant isolation.
  • Own fleet health: monitoring, alerting, and rapid triage of hardware and network faults.
  • Onboard engineers/researchers and assist with workloads when infrastructure is the bottleneck.
  • Partner with product teams to translate compute needs into platform capability.

Skills

GPU compute at scale
Kubernetes
Automation pipelines
Physical datacenter work

Tools

H200/B200/B300 GPUs
NVLink/InfiniBand/RoCE
VAST, DDN, Weka, Lustre
Kubernetes Run:AI

Job description

Anduril Industries is seeking a Senior AI Infrastructure Engineer to own the stability and scalability of large-scale GPU training. You will drive automation and optimize Kubernetes/Run:AI for research and product teams.

This role requires 10+ years in HPC/datacenter infra, hands-on GPU hardware experience, and eligibility for a U.S. Top Secret clearance. You’ll build self-healing systems and improve interconnects to support massive-scale compute.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI Infra Engineer: Large-Scale GPU & HPC
Senior AI Infra Engineer: Large-Scale GPU & HPC

Anduril Industries • Costa Mesa (CA)

On-site
USD 166,000 - 220,000
Equity grants
Benefits package
Senior GPU Infra Engineer for Distributed AI
Senior GPU Infra Engineer for Distributed AI

Andromeda Cluster • San Francisco (CA)

Hybrid
USD 180,000 - 240,000
Senior AI Infrastructure Engineer — Scale GPU Clusters Remote
Senior AI Infrastructure Engineer — Scale GPU Clusters Remote

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 280,000 - 420,000
Equity options
Health, vision, dental benefits
Unlimited PTO
+2
Senior AI Infrastructure Engineer, Physical Infrastructure
Senior AI Infrastructure Engineer, Physical Infrastructure

AI Chopping Block • Costa Mesa (CA), Northern (KY)

Hybrid
USD 166,000 - 220,000
AI Infrastructure Architect — Scalable GPU Compute
AI Infrastructure Architect — Scalable GPU Compute

EngineersOfAI • Sunnyvale (CA)

On-site
USD 150,000 - 200,000
Senior GPU Compute Solutions Architect
Senior GPU Compute Solutions Architect

Computacenter AG & Co. oHG • Northern (KY)

Hybrid
USD 190,000 - 230,000
Staff Compute Infra Engineer - GPU & AI Systems
Staff Compute Infra Engineer - GPU & AI Systems

xAI • Palo Alto (CA)

On-site
USD 180,000 - 440,000
Senior AI Infrastructure Engineer — Scale GPU Clusters
Senior AI Infrastructure Engineer — Scale GPU Clusters

AI Breaking Wire • San Francisco (CA)

On-site
USD 280,000 - 400,000
Equity
Medical, dental, and vision benefits
Unlimited PTO
+2
Senior AI Infrastructure Engineer, Physical Infrastructure
Senior AI Infrastructure Engineer, Physical Infrastructure

Anduril Industries • Costa Mesa (CA)

On-site
USD 166,000 - 220,000
Equity grants
Benefits package
Head of AI Data Center Infrastructure Platforms and Software
Head of AI Data Center Infrastructure Platforms and Software

Summit Group Solutions, LLC • United States

On-site
USD 150,000 - 350,000