Principal Infrastructure Engineer, AI Cluster Performance & Validation

Nscale

New York, San Francisco, Seattle (NY, CA, WA)

On-site

USD 180,000 - 240,000

Full time

9 days ago
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Nscale seeks a Principal Infrastructure Engineer to own AI cluster performance and validation across thousands of GPUs. You’ll define healthy-at-scale criteria, set the architecture, and partner with senior leaders to production-readiness goals.

You’ll run real AI workloads at scale, diagnose performance regressions, design scalable validation systems, and optimize fabric, storage, and schedulers to deliver reliable, cost-efficient AI clusters.

Qualifications

  • Education: Bachelor's or higher in Computer Science, Engineering or related field.
  • Experience: 10+ years in large-scale compute infrastructure with cross-team technical direction.
  • AI Workload Expertise: Hands-on experience with real AI compute jobs at scale and distributed training strategies.
  • Cluster Validation: Experience validating and accepting large clusters (thousands of GPUs) for performance and reliability.
  • Performance Debugging: Ability to diagnose distributed performance problems using NCCL, Nsight, PyTorch Profiler, perf, etc.
  • Networking: Deep understanding of InfiniBand/RoCE, RDMA, GPUDirect and related networking.
  • Systems & Programming: Strong Linux, Python, C/C++ or Go, and IaC tooling.
  • Schedulers: Experience with SLURM and/or Kubernetes at scale.

Responsibilities

  • Define healthy-at-scale criteria and production readiness for thousands of GPUs.
  • Run real AI workloads to validate cluster behavior under genuine load.
  • Diagnose large-scale failures and performance regressions across the full stack.
  • Design and build validation and burn-in systems for scalable qualification.
  • Drive end-to-end cluster optimization of fabric, storage, and schedulers.
  • Collaborate with cross-functional teams to translate needs into durable solutions.
  • Build production-grade Python tooling for triage and telemetry.
  • Establish engineering standards for reliability and observability.

Skills

AI workloads
GPU clusters
PyTorch
Megatron-LM
DeepSpeed
SLURM
Kubernetes
Python
C/C++

Education

Bachelor's degree or higher in CS/CE or equivalent
Master's degree or PhD (preferred)

Tools

NVIDIA Nsight
Perf
Prometheus
Grafana
OpenTelemetry
Terraform
Ansible

Job description

Principal Infrastructure Engineer, AI Cluster Performance & Validation

Houston; New York; San Francisco; Seattle

Overview

As a Principal Infrastructure Engineer, AI Cluster Performance & Validation, you will be a critical member of the AI Infrastructure Operations team, responsible for ensuring the acceptance, performance, and scalability of our cutting-edge AI and High-Performance Computing (HPC) environments. Leveraging software engineering and testing principles, you will focus on building and maintaining the control plane, tooling, and automation that supports performance and validation testing of large-scale AI clusters. Your work will directly translate into higher system availability, compute optimization and reduced operational costs.

Key Responsibilities
  • Own the technical definition of "healthy at scale." Set the architecture, roadmap, and acceptance criteria by which multi-thousand-GPU clusters are declared production-ready, and establish the performance bar (collective bandwidth, job goodput, model FLOPs utilization) that every cluster must clear before and after customer handover. Establish technology and product direction in collaboration with other tech leads, managers, and senior leadership.
  • Run and instrument real AI workloads as a diagnostic instrument. Stand up and execute distributed training and inference jobs — open-source and customer-representative models — across thousands of accelerators to validate cluster behavior under genuine load rather than synthetic proxies alone, and translate what those runs reveal into fleet-wide fixes.
  • Lead deep diagnosis of large-scale cluster failures and performance regressions, isolating root cause across the full stack: GPU and NIC firmware, PCIe/NVLink topology and NUMA placement, InfiniBand/RoCE fabric health, congestion control and routing, storage and data-loader throughput, scheduler placement, and framework/communication-library behavior. Serve as the final escalation point for the hardest slow-job and stalled-job investigations.
  • Design and build the validation and burn-in systems that qualify nodes, racks, and full pods at scale — NCCL/RCCL collective sweeps, HPL/HPCG and MLPerf-style benchmarks, thermal and power soak tests, straggler and flapping-link detection — and automate them so that qualification is a repeatable pipeline, not a manual campaign.
  • Drive cluster optimization end to end, tuning fabric configuration (adaptive routing, QoS and congestion control, SHARP in-network reduction, rail and topology-aware placement), collective communication libraries and algorithm selection, GPUDirect RDMA and storage paths, and host-level settings (huge pages, IRQ affinity, CPU governors, MIG and driver configuration) to convert raw hardware into delivered throughput.
  • Partner with Infrastructure, Platform, SRE, and customer-facing teams to translate operational and customer performance needs into durable engineering solutions, and to feed diagnostic signal back into provisioning, remediation, and capacity workflows.
  • Build production-grade Python systems and performance tooling for automated triage, telemetry correlation, and regression detection, leveraging AI tools to accelerate delivery. Assess impact to the team's software and validation stack from new hardware product programs, and explore AI-driven process improvement and automation.
  • Establish engineering standards for reliability, observability, benchmarking methodology, and operational excellence across all services, and raise the diagnostic capability of the wider organization through mentorship, runbooks, and post-incident technical write-ups.
Required Qualifications
  • Education: Bachelor's or higher degree in Computer Science, Computer Engineering, relevant technical field, or equivalent practical experience.
  • Experience: 10+ years of relevant experience building, operating, or debugging large-scale compute infrastructure, including significant time at staff or principal level owning cross-team technical direction.
  • AI Workload Expertise: Hands-on experience running real AI compute jobs at scale — pre-training, fine-tuning, or large-scale inference of open-source or proprietary models — including practical familiarity with distributed training strategies (data, tensor, pipeline, and expert parallelism) and frameworks such as PyTorch, Megatron-LM, DeepSpeed, or equivalent.
  • Cluster Validation: Demonstrated experience validating and accepting large clusters (thousands of GPUs) for performance and reliability, with a working command of benchmark methodology and the ability to defend a number to both engineers and customers.
  • Performance Debugging: Proven ability to diagnose distributed performance problems — stragglers, collective stalls, link flaps, thermal throttling, silent data corruption, ECC and Xid errors, noisy-neighbor and storage-bound bottlenecks — using tools such as NCCL debug tracing, Nsight Systems/Compute, PyTorch Profiler, perf, and fabric telemetry.
  • Networking: Deep understanding of high-performance fabrics — InfiniBand and/or RoCEv2, RDMA, GPUDirect, adaptive routing, congestion control, and rail-optimized topologies — and of networking fundamentals (TCP/IP, BGP).
  • Systems & Programming: Deep Linux systems expertise (kernel tunables, NUMA, PCIe, IRQ and memory behavior) and strong production Python, plus experience with C/C++ or Go and with configuration management tooling (e.g., Ansible, Terraform).
  • Schedulers: Experience operating and debugging AI workloads under SLURM and/or Kubernetes at scale.
Preferred Qualifications
  • Master's degree or PhD in Engineering, Computer Science, or a related technical field.
  • Experience bringing up and qualifying a greenfield GPU supercluster from first rack to production traffic, including firmware, driver, and topology standardization across a heterogeneous fleet.
  • Direct experience with NVIDIA GPU platforms (H200/GB200/GB300-class), NVLink and NVSwitch domains, DCGM, SHARP, UFM, and the NVIDIA software stack; or equivalent depth on AMD Instinct and ROCm/RCCL.
  • Published or presented benchmark, scaling, or post-mortem work — MLPerf submissions, scaling studies, or public technical write-ups on large-cluster behavior.
  • Experience with advanced observability and monitoring systems (Prometheus, Grafana, OpenTelemetry) applied to high-cardinality GPU and fabric telemetry, including automated anomaly and regression detection.
  • Experience with high-throughput parallel storage (Lustre, GPFS, WEKA, VAST) and with diagnosing data-pipeline-bound training jobs.
  • Familiarity with cloud-native technologies (Kubernetes, Docker), infrastructure-as-code principles, and integration with infrastructure tooling such as DCIMs, NetBox, and bare metal APIs (MAAS, Ironic, IPMI, Redfish).
  • Demonstrated ability to integrate AI tools to optimize/redesign workflows and drive measurable impact (e.g., efficiency gains, quality improvements).
  • Familiarity with SLOs/metrics measurement and logs/telemetry/metrics integration with tools for enhanced operator experience.

For information on how Nscale handles candidate personal data, please see our Employee & Candidate Privacy Notice:Here.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Principal Infrastructure Engineer, AI Cluster Performance & Validation
Principal Infrastructure Engineer, AI Cluster Performance & Validation

Socket.dev • Houston (TX)

On-site
USD 200,000 - 260,000
Cluster Engineer
Cluster Engineer

STN Inc • San Francisco (CA)

On-site
USD 180,000 - 240,000
Head of AI Data Center Infrastructure Platforms and Software
Head of AI Data Center Infrastructure Platforms and Software

Summit Group Solutions, LLC • United States

On-site
USD 150,000 - 350,000
Senior Software Engineer, Infrastructure Software for AI (Centralized AI Data Centers & Distrib[...]
Senior Software Engineer, Infrastructure Software for AI (Centralized AI Data Centers & Distrib[...]

Intelliswift - An LTTS Company • Sunnyvale (CA)

On-site
USD 120,000 - 150,000
Competitive salary
Health insurance
Flexible work hours
Cluster Design
Cluster Design

Blue Signal Search • San Francisco (CA)

On-site
USD 150,000 - 230,000
Infrastructure Software Engineer, Fleet & Automation New Houston; New York; San Francisco; Seattle
Infrastructure Software Engineer, Fleet & Automation New Houston; New York; San Francisco; Seattle

Nscale • New York (NY)

On-site
USD 140,000 - 210,000
Competitive package
Equity
Growth opportunities
Senior HPC AI Cluster Engineer
Senior HPC AI Cluster Engineer

NVIDIA • California (MO)

On-site
USD 176,000 - 334,000
Senior Network Architect – AI Infrastructure
Senior Network Architect – AI Infrastructure

Nscale • United States

On-site
USD 120,000 - 150,000
Senior HPC AI Cluster Engineer
Senior HPC AI Cluster Engineer

NVIDIA AI • Santa Clara (CA)

On-site
USD 176,000 - 334,000
Equity
Benefits
Senior Network Architect – AI Infrastructure
Senior Network Architect – AI Infrastructure

Nscale • Seattle (WA)

On-site
USD 120,000 - 160,000