AI and ML Infra Software Engineer, GPU Clusters - New College Grad 2026

NVIDIA AI

Redmond (WA)

On-site

USD 120,000 - 180,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Competitive salaries
Comprehensive benefits package
Equity

Job summary

NVIDIA AI in Redmond seeks a recent graduate with an MS/PhD in CS to collaborate with AI/ML research teams, identify infrastructure gaps, and implement scalable solutions on GPU clusters.

You will monitor performance, optimize for high availability, and ensure researchers have efficient resources using PyTorch, Kubernetes, Slurm, and Docker.

Role focuses on distributed training, data processing, and model inference across HPC workloads with equity and comprehensive benefits.

Qualifications

  • MS/PhD in CS or equivalent with HPC/accelerated computing exposure.
  • Experience with distributed training and GPU-accelerated frameworks.
  • Proficiency in Python; familiarity with Go is a plus.

Responsibilities

  • Collaborate with AI/ML research teams to identify gaps and implement scalable GPU infra.
  • Monitor and optimize infrastructure performance for high availability.
  • Ensure efficient resource utilization for researchers and their workloads.

Skills

AI/ML Infrastructure
HPC Workloads
GPU Computing
PyTorch
Kubernetes
Slurm
Docker
Python
Go
Bash
Distributed Training
Infiniband
Cloud Computing
Data Processing
Model Inference
Parallel Computing

Education

MS in Computer Science
PhD in Computer Science
Equivalent experience

Tools

Kubernetes
Slurm
Docker

Job description

Collaborate with AI/ML research teams to identify infrastructure gaps and implement scalable solutions on GPU clusters. Monitor and optimize infrastructure performance to ensure high availability and efficient resource utilization for researchers.

Requirements:

Requires a recent graduate with a MS, PhD, or equivalent in Computer Science with experience in HPC and accelerated computing. Proficiency in Python, Go, and distributed training frameworks like PyTorch or JAX is essential.

Key Skills:

AI/ML Infrastructure, HPC Workloads, GPU Computing, PyTorch, Kubernetes, Slurm, Docker, Python, Go, Bash, Distributed Training, Infiniband, Cloud Computing, Data Processing, Model Inference, Parallel Computing

Benefits:
  • Competitive salaries
  • Comprehensive benefits package
  • Equity
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI/ML Infra Engineer — GPU Clusters & HPC
AI/ML Infra Engineer — GPU Clusters & HPC

NVIDIA AI • Redmond (WA)

On-site
USD 120,000 - 180,000
Competitive salaries
Comprehensive benefits package
Equity
AI/ML Infra Engineer - Hosting
AI/ML Infra Engineer - Hosting

Hamilton Barnes Associates Limited • San Francisco (CA)

On-site
USD 225,000 - 275,000
Stock options
Software Engineer, AI Infra
Software Engineer, AI Infra

Makers Fund • Palo Alto (CA)

On-site
USD 180,000 - 240,000
Equity
Health benefits
Monthly stipends
+1
AI and ML HPC Cluster Engineer, AI and ML HPC Cluster Engineer
AI and ML HPC Cluster Engineer, AI and ML HPC Cluster Engineer

NVIDIA • Colorado

On-site
USD 124,000 - 196,000
AI Kernel / Cluster Engineer
AI Kernel / Cluster Engineer

Blue Signal Search • Santa Clara (CA)

On-site
USD 150,000 - 210,000
Cluster Engineer
Cluster Engineer

STN Inc • San Francisco (CA)

On-site
USD 180,000 - 240,000
Cluster Design
Cluster Design

Blue Signal Search • San Francisco (CA)

On-site
USD 150,000 - 230,000
Head of AI Data Center Infrastructure Platforms and Software
Head of AI Data Center Infrastructure Platforms and Software

Summit Group Solutions, LLC • United States

On-site
USD 150,000 - 350,000
Senior HPC AI Cluster Engineer
Senior HPC AI Cluster Engineer

NVIDIA • California (MO)

On-site
USD 176,000 - 334,000
AI/ML Engineer – (Next-Generation AI Platforms & Workloads)
AI/ML Engineer – (Next-Generation AI Platforms & Workloads)

VeeAR Projects Inc. • Sunnyvale (CA)

On-site
USD 140,000 - 210,000