Senior AI GPU Infrastructure Architect

ECLARO

Costa Mesa (CA)

On-site

USD 166,000 - 220,000

Full time

9 days ago
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

ECLARO seeks a Senior AI Infrastructure Engineer in Costa Mesa, CA to lead GPU-scale training infrastructure. You will own the stability of GPU clusters, automate resilience, and optimize NCCL, networking, and scheduling tools to support ML research across multiple teams.

The role demands hands-on hardware management, strong automation, and collaboration with product groups to forecast compute needs and deliver scalable platform capabilities.

Qualifications

  • 10+ years in hands-on infrastructure, HPC, or datacenter engineering.
  • Experience with H200/B200/B300 GPUs and firmware/driver management.
  • Kubernetes required; Run:AI or similar GPU scheduling experience preferred.
  • Able to lift/move 50+ lbs and perform physical datacenter work.
  • Eligible to obtain and maintain an active U.S. Top Secret clearance.

Responsibilities

  • Rack, stack, cable, and bring up GPU compute nodes and validate hardware.
  • Build and tune interconnects (NVLink, InfiniBand, RoCE, Spectrum-X) across hundreds of GPUs.
  • Integrate high-performance storage (VAST, DDN, Weka) for large datasets.
  • Automate cluster deployment end-to-end with infrastructure as code.
  • Operate and extend Kubernetes/Run:AI for GPU scheduling and multi-tenant workloads.
  • Own fleet health: monitoring, alerting, and rapid triage of hardware and network faults.
  • Onboard engineers/researchers to the platform and assist with workload optimization.
  • Partner with product teams to translate compute needs into platform capabilities.

Skills

GPU Compute
Kubernetes
Automation
Distributed Training

Tools

Run:AI
NCCL
InfiniBand
DCGM
Lustre

Job description

ECLARO seeks a Senior AI Infrastructure Engineer in Costa Mesa, CA to lead GPU-scale training infrastructure. You will own the stability of GPU clusters, automate resilience, and optimize NCCL, networking, and scheduling tools to support ML research across multiple teams.

The role demands hands-on hardware management, strong automation, and collaboration with product groups to forecast compute needs and deliver scalable platform capabilities.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Sr. AI Infrastructure Engineer
Sr. AI Infrastructure Engineer

ECLARO • Costa Mesa (CA)

On-site
USD 166,000 - 220,000
Senior AI Infrastructure Lead: GPU Clusters & LLMs
Senior AI Infrastructure Lead: GPU Clusters & LLMs

Cadence Design Systems • San Jose (CA)

On-site
USD 137,000 - 254,000
Senior AI GPU Cluster Architect
Senior AI GPU Cluster Architect

STN Inc • San Francisco (CA)

On-site
USD 180,000 - 240,000
Senior AI Infrastructure Engineer - GPU & Kubernetes
Senior AI Infrastructure Engineer - GPU & Kubernetes

HCL Technologies Limited • California (MO)

On-site
USD 120,000 - 180,000
401(k) retirement plan
Paid time off (PTO)
Paid holidays
+1
Senior AI Infrastructure Engineer — Scale GPU Clusters Remote
Senior AI Infrastructure Engineer — Scale GPU Clusters Remote

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 280,000 - 420,000
Equity options
Health, vision, dental benefits
Unlimited PTO
+2
Remote AI Research Clusters Engineer - ML Infra & GPU
Remote AI Research Clusters Engineer - ML Infra & GPU

NEPSE Trading • Northern (KY)

Hybrid
USD 124,000 - 196,000
Senior Platform Systems Engineer – AI & Data Center Infra
Senior Platform Systems Engineer – AI & Data Center Infra

ECLARO • San Diego (CA)

On-site
USD 180,000 - 240,000
401k Retirement Savings Plan
Commuter Check Pretax Benefits
Medical, Dental & Vision Insurance
Lead AI Infrastructure Architect for Large-Scale GPU Clusters
Lead AI Infrastructure Architect for Large-Scale GPU Clusters

NVIDIA • Santa Clara (CA)

On-site
USD 184,000 - 356,500
Equity and benefits
Senior AI Infra Engineer — Large-Scale GPU Cloud Equity
Senior AI Infra Engineer — Large-Scale GPU Cloud Equity

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Senior AI Factory Deployment Architect (Multi-GPU)
Senior AI Factory Deployment Architect (Multi-GPU)

NVIDIA • Virginia (MN)

On-site
USD 148,000 - 235,750