Senior AI Infra Engineer: GPU Scale & Resilience (Hybrid)

Calance

Costa Mesa (CA)

Hybrid

USD 180,000 - 240,000

Full time

3 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Calance seeks a Senior AI Infrastructure Engineer to own the scalable GPU training pipeline, from bring-up to operation. You will design fault-tolerant, self-healing GPU clusters, tune high-speed interconnects, and automate deployment with IaC.

You will enhance NCCL, Run:AI, and Ray scheduling, ensuring high availability for ML workloads across teams. You will lead hardware bring-up, firmware management, and fabric optimization while maintaining fault isolation and observability in a hybrid

Qualifications

  • Hands-on experience in infrastructure/HPC/datacenter engineering for GPU compute at scale.
  • Experience with high-performance interconnects (NVLink, InfiniBand, RoCE) in large GPU clusters.
  • Experience with Kubernetes and GPU scheduling (Run:AI or similar) is highly preferred.

Responsibilities

  • Rack, stack, cabling, and bring up GPU compute nodes including topology, power, cooling, BIOS, and burn-in validation.
  • Build and tune interconnect fabrics (NVLink, InfiniBand, RoCE, Spectrum-X) for large GPU clusters.
  • Automate cluster deployment/configuration end-to-end with infrastructure as code; manage firmware/driver.
  • Operate and extend Kubernetes/Run:AI for GPU scheduling and multi-tenant isolation.
  • Own fleet health: monitoring, alerting, and rapid triage of hardware/network faults.
  • Onboard researchers/engineers; assist debugging/training workloads when infrastructure is the bottleneck.
  • Collaborate with product teams to translate compute needs into platform capabilities.

Skills

Automation
Cluster management
Hardware fault diagnosis
High performance computing

Tools

Kubernetes
Run:AI
Ray
NCCL
NVLink
InfiniBand
VAST/NDN/Weka

Job description

Calance seeks a Senior AI Infrastructure Engineer to own the scalable GPU training pipeline, from bring-up to operation. You will design fault-tolerant, self-healing GPU clusters, tune high-speed interconnects, and automate deployment with IaC.

You will enhance NCCL, Run:AI, and Ray scheduling, ensuring high availability for ML workloads across teams. You will lead hardware bring-up, firmware management, and fabric optimization while maintaining fault isolation and observability in a hybrid

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI Infrastructure Engineer — Scale GPU Clusters Remote
Senior AI Infrastructure Engineer — Scale GPU Clusters Remote

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 280,000 - 420,000
Equity options
Health, vision, dental benefits
Unlimited PTO
+2
Senior AI Compute Infra Engineer (Hybrid)
Senior AI Compute Infra Engineer (Hybrid)

Arm Limited • Seattle (WA)

Hybrid
USD 209,000 - 283,000
Relocation package
Recruitment accommodations
Senior AI Infrastructure Engineer — Scale GPU Clusters
Senior AI Infrastructure Engineer — Scale GPU Clusters

AI Breaking Wire • San Francisco (CA)

On-site
USD 280,000 - 400,000
Equity
Medical, dental, and vision benefits
Unlimited PTO
+2
Senior ML Infra Engineer - Scale GPU Clusters, Remote
Senior ML Infra Engineer - Scale GPU Clusters, Remote

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 320,000 - 500,000
Equity
Medical/Dental/Vision coverage
Unlimited PTO
+1
Senior GPU Infra Architect for Scalable AI Compute
Senior GPU Infra Architect for Scalable AI Compute

AI Chopping Block • Costa Mesa (CA), Northern (KY)

Hybrid
USD 166,000 - 220,000
Artificial Intelligence Engineer
Artificial Intelligence Engineer

Calance • Costa Mesa (CA)

Hybrid
USD 180,000 - 240,000
Lead AI Infrastructure Architect for Large-Scale GPU Clusters
Lead AI Infrastructure Architect for Large-Scale GPU Clusters

NVIDIA • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Equity and benefits
Senior AI Infrastructure Engineer – Scale & Equity
Senior AI Infrastructure Engineer – Scale & Equity

Nvidia Corporation • Santa Clara (CA)

On-site
USD 184,000 - 356,500
Equity
Benefits
Senior AI Infrastructure Engineer | Scale GPU Clusters
Senior AI Infrastructure Engineer | Scale GPU Clusters

Fuel Talent LLC • Seattle (WA)

Hybrid
USD 126,000 - 189,000
Senior AI Infrastructure Lead: GPU Clusters & LLMs
Senior AI Infrastructure Lead: GPU Clusters & LLMs

Cadence Design Systems • San Jose (CA)

On-site
USD 137,000 - 254,000