Sr. AI Infrastructure Engineer

ECLARO

Costa Mesa (CA)

On-site

USD 166,000 - 220,000

Full time

8 days ago
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

ECLARO seeks a Senior AI Infrastructure Engineer in Costa Mesa, CA to lead GPU-scale training infrastructure. You will own the stability of GPU clusters, automate resilience, and optimize NCCL, networking, and scheduling tools to support ML research across multiple teams.

The role demands hands-on hardware management, strong automation, and collaboration with product groups to forecast compute needs and deliver scalable platform capabilities.

Qualifications

  • 10+ years in hands-on infrastructure, HPC, or datacenter engineering.
  • Experience with H200/B200/B300 GPUs and firmware/driver management.
  • Kubernetes required; Run:AI or similar GPU scheduling experience preferred.
  • Able to lift/move 50+ lbs and perform physical datacenter work.
  • Eligible to obtain and maintain an active U.S. Top Secret clearance.

Responsibilities

  • Rack, stack, cable, and bring up GPU compute nodes and validate hardware.
  • Build and tune interconnects (NVLink, InfiniBand, RoCE, Spectrum-X) across hundreds of GPUs.
  • Integrate high-performance storage (VAST, DDN, Weka) for large datasets.
  • Automate cluster deployment end-to-end with infrastructure as code.
  • Operate and extend Kubernetes/Run:AI for GPU scheduling and multi-tenant workloads.
  • Own fleet health: monitoring, alerting, and rapid triage of hardware and network faults.
  • Onboard engineers/researchers to the platform and assist with workload optimization.
  • Partner with product teams to translate compute needs into platform capabilities.

Skills

GPU Compute
Kubernetes
Automation
Distributed Training

Tools

Run:AI
NCCL
InfiniBand
DCGM
Lustre

Job description

Job Number: 26-01607 Use your skills where innovative technology solutions begin. ECLARO is looking for a

Costa Mesa, CA. ECLARO’s client is a leading technology solutions provider, collaborating with customers to manage their needs and achieve success in their business goals. If you’re up to the challenge, then take a chance at this rewarding opportunity!

Senior AI Infrastructure Engineer

Job Number: 26-01607 Use your skills where innovative technology solutions begin. ECLARO is looking for a Senior AI Infrastructure Engineer for our client in Costa Mesa, CA. ECLARO’s client is a leading technology solutions provider, collaborating with customers to manage their needs and achieve success in their business goals. If you’re up to the challenge, then take a chance at this rewarding opportunity!

Position Overview
  • A Senior AI Infrastructure Engineer to lead the vision, execution, and long-term stability of how Company trains with GPUs at scale.
  • In this role, you will take absolute ownership of cluster robustness, ensuring our high-performance GPU systems are highly available, fault-tolerant, and resilient for ML platform and research teams company-wide.
  • This is a highly hands-on role where your primary focus is logical stability and automated resilience-building self-healing mechanisms to proactively detect and isolate hardware faults, tuning NCCL and high-speed networking, and optimizing Kubernetes, Run:AI, and Ray scheduling.
  • By replacing manual triage with automated deployment tooling and deep observability, you will ensure our massive-scale training infrastructure runs seamlessly and scales without linear headcount growth.
Responsibilities
  • Rack, stack, cable, and bring up GPU compute (H200/B200/B300, NVL72) including physical topology, power, cooling, firmware/BIOS, and burn in validation.
  • Build and tune the interconnect fabric (NVLink, InfiniBand, RoCE, Spectrum-X) connecting hundreds of GPUs into low latency training and inference clusters.
  • Integrate high performance parallel storage (VAST, DDN, Weka) to sustain the throughput demanded by distributed training and terabyte scale multi modal datasets across Company's programs.
  • Automate cluster deployment and configuration end to end, including infrastructure as code for bring up, firmware/driver management, and fabric config, so new capacity comes online with minimal manual work.
  • Operate and extend our Kubernetes/Run:AI environment for GPU scheduling, quota management, and multi tenant workload isolation across research and engineering teams company wide.
  • Own fleet health: monitoring, alerting, and rapid triage of hardware and network faults (bad transceivers, GPU Xid errors, NCCL/collective failures, RoCE congestion).
  • Onboard engineers and researchers onto the platform and act as their escalation point, working directly alongside them to debug, train and optimize their workloads whenever infrastructure, not the model, is the bottleneck.
  • Partner with product facing teams across Company to understand emerging compute needs and translate them into platform capability.
Required Qualifications
  • 10+ years in a hands on infrastructure, HPC, or datacenter engineering role supporting GPU compute at scale.
  • Hands on experience with H200/B200/B300 (or comparable) GPU systems: bring up, cabling, firmware/driver management.
  • Experience with high performance interconnects (NVLink, InfiniBand, RoCE, Spectrum-X) in clusters of hundreds of GPUs.
  • Experience with high performance parallel storage (VAST, DDN, Weka, Lustre, or similar).
  • Kubernetes required; Run:AI or similar GPU scheduling/orchestration experience strongly preferred.
  • Strong automation background. You build repeatable, automated deployment pipelines rather than manual processes.
  • Able to lift/move 50+ lbs. and perform physical datacenter work (rack/stack/cable/troubleshoot).
  • Eligible to obtain and maintain an active U.S. Top Secret clearance.
Preferred Qualifications
  • Experience with NVIDIA NVL72 rack scale systems.
  • Experience supporting LLM token serving/inference infrastructure alongside training clusters.
  • Network fabric tuning experience (congestion control, adaptive routing, QoS) for RoCE/InfiniBand at scale.
  • Familiarity with GPU/network observability tooling (DCGM, fabric telemetry) and automated fault detection.
  • Experience supporting infrastructure as a shared platform serving multiple internal customer teams with differing requirements.
Salary

$166000.00-$220000.00/Year.

Equal Opportunity Employer

ECLARO values diversity and does not discriminate based on Race, Color, Religion, Sex, Sexual Orientation, National Origin, Age, Genetic Information, Disability, Protected Veteran Status, or any other legally protected group status, in compliance with all applicable laws.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Head of AI Data Center Infrastructure Platforms and Software
Head of AI Data Center Infrastructure Platforms and Software

Summit Group Solutions, LLC • United States

On-site
USD 150,000 - 350,000
Senior AI GPU Infrastructure Architect
Senior AI GPU Infrastructure Architect

ECLARO • Costa Mesa (CA)

On-site
USD 166,000 - 220,000
Senior AI Infrastructure Engineer, Physical Infrastructure
Senior AI Infrastructure Engineer, Physical Infrastructure

AI Chopping Block • Costa Mesa (CA), Northern (KY)

Hybrid
USD 166,000 - 220,000
Senior HPC AI Cluster Engineer
Senior HPC AI Cluster Engineer

NVIDIA • United States

On-site
USD 176,000 - 334,000
Equity
Benefits
Senior AI Performance and Efficiency Engineer
Senior AI Performance and Efficiency Engineer

NVIDIA • California (MO)

On-site
USD 152,000 - 288,000
Equity
Competitive benefits
Principal Infrastructure Engineer, AI Cluster Performance & Validation
Principal Infrastructure Engineer, AI Cluster Performance & Validation

Nscale • New York (NY), San Francisco (CA), Seattle (WA)

On-site
USD 180,000 - 240,000
Senior HPC Cluster Engineer
Senior HPC Cluster Engineer

NVIDIA • Santa Clara (CA)

On-site
USD 152,000 - 241,500
Equity
Benefits
Infrastructure Engineer
Infrastructure Engineer

HCL Technologies Limited • California (MO)

On-site
USD 120,000 - 180,000
401(k) retirement plan
Paid time off (PTO)
Paid holidays
+1
Senior Solution Architect, AI Infrastructure
Senior Solution Architect, AI Infrastructure

NVIDIA • Washington

On-site
USD 184,000 - 357,000
Equity compensation
Benefits
Senior Solutions Architect, Generative AI
Senior Solutions Architect, Generative AI

NVIDIA • California (MO)

Hybrid
USD 184,000 - 357,000