Senior AI Cluster Performance & Validation Engineer

Nscale

New York, San Francisco, Seattle (NY, CA, WA)

On-site

USD 180,000 - 240,000

Full time

9 days ago
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Nscale seeks a Principal Infrastructure Engineer to own AI cluster performance and validation across thousands of GPUs. You’ll define healthy-at-scale criteria, set the architecture, and partner with senior leaders to production-readiness goals.

You’ll run real AI workloads at scale, diagnose performance regressions, design scalable validation systems, and optimize fabric, storage, and schedulers to deliver reliable, cost-efficient AI clusters.

Qualifications

  • Education: Bachelor's or higher in Computer Science, Engineering or related field.
  • Experience: 10+ years in large-scale compute infrastructure with cross-team technical direction.
  • AI Workload Expertise: Hands-on experience with real AI compute jobs at scale and distributed training strategies.
  • Cluster Validation: Experience validating and accepting large clusters (thousands of GPUs) for performance and reliability.
  • Performance Debugging: Ability to diagnose distributed performance problems using NCCL, Nsight, PyTorch Profiler, perf, etc.
  • Networking: Deep understanding of InfiniBand/RoCE, RDMA, GPUDirect and related networking.
  • Systems & Programming: Strong Linux, Python, C/C++ or Go, and IaC tooling.
  • Schedulers: Experience with SLURM and/or Kubernetes at scale.

Responsibilities

  • Define healthy-at-scale criteria and production readiness for thousands of GPUs.
  • Run real AI workloads to validate cluster behavior under genuine load.
  • Diagnose large-scale failures and performance regressions across the full stack.
  • Design and build validation and burn-in systems for scalable qualification.
  • Drive end-to-end cluster optimization of fabric, storage, and schedulers.
  • Collaborate with cross-functional teams to translate needs into durable solutions.
  • Build production-grade Python tooling for triage and telemetry.
  • Establish engineering standards for reliability and observability.

Skills

AI workloads
GPU clusters
PyTorch
Megatron-LM
DeepSpeed
SLURM
Kubernetes
Python
C/C++

Education

Bachelor's degree or higher in CS/CE or equivalent
Master's degree or PhD (preferred)

Tools

NVIDIA Nsight
Perf
Prometheus
Grafana
OpenTelemetry
Terraform
Ansible

Job description

Nscale seeks a Principal Infrastructure Engineer to own AI cluster performance and validation across thousands of GPUs. You’ll define healthy-at-scale criteria, set the architecture, and partner with senior leaders to production-readiness goals.

You’ll run real AI workloads at scale, diagnose performance regressions, design scalable validation systems, and optimize fabric, storage, and schedulers to deliver reliable, cost-efficient AI clusters.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Principal AI Cluster Performance & Validation Architect
Principal AI Cluster Performance & Validation Architect

Socket.dev • Houston (TX)

On-site
USD 200,000 - 260,000
Principal Infrastructure Engineer, AI Cluster Performance & Validation
Principal Infrastructure Engineer, AI Cluster Performance & Validation

Nscale • New York (NY), San Francisco (CA), Seattle (WA)

On-site
USD 180,000 - 240,000
Principal Infrastructure Engineer, AI Cluster Performance & Validation
Principal Infrastructure Engineer, AI Cluster Performance & Validation

Socket.dev • Houston (TX)

On-site
USD 200,000 - 260,000
Senior AI Cluster Performance Validation Engineer
Senior AI Cluster Performance Validation Engineer

Advanced Micro Devices • Austin (TX), Northern (KY)

Hybrid
USD 180,000 - 230,000
Senior AI Performance & Efficiency Engineer - Equity Eligible
Senior AI Performance & Efficiency Engineer - Equity Eligible

NVIDIA • California (MO)

On-site
USD 152,000 - 288,000
Equity
Competitive benefits
Senior AI Infrastructure Engineer — Scale GPU Clusters
Senior AI Infrastructure Engineer — Scale GPU Clusters

AI Breaking Wire • San Francisco (CA)

On-site
USD 280,000 - 400,000
Equity
Medical, dental, and vision benefits
Unlimited PTO
+2
Senior Observability Platform Engineer – AI GPU Scale
Senior Observability Platform Engineer – AI GPU Scale

Nscale • United States

On-site
USD 160,000 - 230,000
Medical, dental, vision insurance
Flexible paid time off (PTO)
Parental leave
+1
Senior AI Observability Platform Engineer
Senior AI Observability Platform Engineer

Socket.dev • United States

On-site
USD 160,000 - 230,000
Senior AI Infrastructure Validation Architect
Senior AI Infrastructure Validation Architect

Mainz Brady Group • Portland (OR)

On-site
USD 150,000 - 190,000
Senior GPU Cluster Performance Validation Engineer
Senior GPU Cluster Performance Validation Engineer

AMD • Austin (TX)

On-site
USD 150,000 - 190,000
AMD Benefits