AI/ML Infra Engineer — GPU Clusters & HPC

NVIDIA AI

Redmond (WA)

On-site

USD 120,000 - 180,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Competitive salaries
Comprehensive benefits package
Equity

Job summary

NVIDIA AI in Redmond seeks a recent graduate with an MS/PhD in CS to collaborate with AI/ML research teams, identify infrastructure gaps, and implement scalable solutions on GPU clusters.

You will monitor performance, optimize for high availability, and ensure researchers have efficient resources using PyTorch, Kubernetes, Slurm, and Docker.

Role focuses on distributed training, data processing, and model inference across HPC workloads with equity and comprehensive benefits.

Qualifications

  • MS/PhD in CS or equivalent with HPC/accelerated computing exposure.
  • Experience with distributed training and GPU-accelerated frameworks.
  • Proficiency in Python; familiarity with Go is a plus.

Responsibilities

  • Collaborate with AI/ML research teams to identify gaps and implement scalable GPU infra.
  • Monitor and optimize infrastructure performance for high availability.
  • Ensure efficient resource utilization for researchers and their workloads.

Skills

AI/ML Infrastructure
HPC Workloads
GPU Computing
PyTorch
Kubernetes
Slurm
Docker
Python
Go
Bash
Distributed Training
Infiniband
Cloud Computing
Data Processing
Model Inference
Parallel Computing

Education

MS in Computer Science
PhD in Computer Science
Equivalent experience

Tools

Kubernetes
Slurm
Docker

Job description

NVIDIA AI in Redmond seeks a recent graduate with an MS/PhD in CS to collaborate with AI/ML research teams, identify infrastructure gaps, and implement scalable solutions on GPU clusters.

You will monitor performance, optimize for high availability, and ensure researchers have efficient resources using PyTorch, Kubernetes, Slurm, and Docker.

Role focuses on distributed training, data processing, and model inference across HPC workloads with equity and comprehensive benefits.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI and ML Infra Software Engineer, GPU Clusters - New College Grad 2026
AI and ML Infra Software Engineer, GPU Clusters - New College Grad 2026

NVIDIA AI • Redmond (WA)

On-site
USD 120,000 - 180,000
Competitive salaries
Comprehensive benefits package
Equity
GPU AI/ML Infra Engineer for HPC Clusters
GPU AI/ML Infra Engineer for HPC Clusters

Jobtailor • California (MO)

On-site
USD 120,000 - 190,000
Senior AI Performance & Efficiency Engineer - Equity Eligible
Senior AI Performance & Efficiency Engineer - Equity Eligible

NVIDIA • California (MO)

On-site
USD 152,000 - 288,000
Equity
Competitive benefits
Hybrid AI HPC Infrastructure Engineer (GPU/ML)
Hybrid AI HPC Infrastructure Engineer (GPU/ML)

Analysis Group, Inc. • Boston (MA)

On-site
USD 150,000 - 170,000
Discretionary annual bonus
Benefits package
Senior AI Infrastructure Engineer — Scale GPU Clusters Remote
Senior AI Infrastructure Engineer — Scale GPU Clusters Remote

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 280,000 - 420,000
Equity options
Health, vision, dental benefits
Unlimited PTO
+2
AI Systems Engineer - Distributed, Multi-GPU (Equity)
AI Systems Engineer - Distributed, Multi-GPU (Equity)

NVIDIA AI • Eugene (OR)

On-site
USD 120,000 - 180,000
Equity
Health Insurance
Senior AI Systems Tools Engineer - GPU Clusters
Senior AI Systems Tools Engineer - GPU Clusters

NVIDIA Gruppe • Santa Clara (CA)

Hybrid
USD 184,000 - 357,000
Cluster Engineer
Cluster Engineer

STN Inc • San Francisco (CA)

On-site
USD 180,000 - 240,000
AI and ML HPC Cluster Engineer, AI and ML HPC Cluster Engineer
AI and ML HPC Cluster Engineer, AI and ML HPC Cluster Engineer

NVIDIA • Colorado

On-site
USD 124,000 - 196,000
AI Kernel / Cluster Engineer
AI Kernel / Cluster Engineer

Blue Signal Search • Santa Clara (CA)

On-site
USD 150,000 - 210,000