Senior GPU Compute Scheduler – Scale AI Clusters

NVIDIA AI

Redmond (WA)

On-site

USD 152,000 - 288,000

Full time

3 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity
Benefits

Job summary

NVIDIA in Redmond is seeking a Scheduling Engineer to design and implement features for GPU compute clusters, enabling large multi-node AI workloads. You will focus on resource fairness, GPU occupancy, resilience, and performance across DL pipelines.

The role requires 5+ years in systems programming with C/C++, Go, Python, Bash, Linux, and experience with container tech (Docker/Singularity/Podman) and batch schedulers (SLURM, K8s). Equity and benefits are offered.

Qualifications

  • Bachelor’s degree in CS, EE or related field or equivalent experience.
  • 5+ years of work experience.
  • Strong knowledge of batch scheduling (SLURM or K8s batch schedulers).
  • Experience in C/C++, Go and scripting (Python, Bash).

Responsibilities

  • Design and develop new scheduling features for GPU clusters.
  • Develop batch workload management and orchestration services.
  • Provide support to staff and end users for batch scheduler issues.
  • Build and improve ecosystem around GPU-accelerated computing.
  • Performance analysis and optimization of deep learning workflows.
  • Develop large-scale automation solutions.
  • Root cause analysis and corrective actions.
  • Identify and fix problems before they occur.

Skills

C/C++
Go
Python
Bash
Linux
Docker/Singularity/Podman
Batch scheduling
Performance tuning
Communication

Education

Bachelor’s degree in Computer Science, Electrical Engineering or related field

Tools

SLURM
Kubernetes batch schedulers (Kueue, Volcano)

Job description

NVIDIA in Redmond is seeking a Scheduling Engineer to design and implement features for GPU compute clusters, enabling large multi-node AI workloads. You will focus on resource fairness, GPU occupancy, resilience, and performance across DL pipelines.

The role requires 5+ years in systems programming with C/C++, Go, Python, Bash, Linux, and experience with container tech (Docker/Singularity/Podman) and batch schedulers (SLURM, K8s). Equity and benefits are offered.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior GPU HPC Scheduler Engineer | Equity Eligible
Senior GPU HPC Scheduler Engineer | Equity Eligible

NVIDIA • Redmond (WA)

On-site
USD 152,000 - 242,000
Senior AI Infrastructure Engineer — Scalable GPU Clusters
Senior AI Infrastructure Engineer — Scalable GPU Clusters

NVIDIA AI • Santa Clara (CA)

On-site
USD 152,000 - 288,000
Senior AI Infrastructure Engineer — GPU Clusters
Senior AI Infrastructure Engineer — GPU Clusters

Nvidia Corporation • Santa Clara (CA)

On-site
USD 152,000 - 288,000
Equity
Benefits
Senior GPU Supercomputer Scheduler Engineer
Senior GPU Supercomputer Scheduler Engineer

NVIDIA AI • Redmond (WA)

On-site
USD 152,000 - 288,000
Equity
Benefits
Senior AI Infrastructure Engineer | Scale GPU Clusters
Senior AI Infrastructure Engineer | Scale GPU Clusters

Fuel Talent LLC • Seattle (WA)

Hybrid
USD 126,000 - 189,000
Senior GPU Supercomputer Scheduler Engineer
Senior GPU Supercomputer Scheduler Engineer

NVIDIA • Redmond (WA)

On-site
USD 152,000 - 242,000
Equity
Benefits
Senior Full-Stack Engineer, AI Infra for GPU Clusters
Senior Full-Stack Engineer, AI Infra for GPU Clusters

NVIDIA • California (MO)

On-site
USD 184,000 - 357,000
Equity
Benefits
Senior HPC AI Cluster Architect — Equity Eligible
Senior HPC AI Cluster Architect — Equity Eligible

NVIDIA Corporation • Santa Clara (CA), Northern (KY)

Hybrid
USD 176,000 - 334,000
Senior GPU Systems Engineer: Scale AI Clusters & HPC
Senior GPU Systems Engineer: Scale AI Clusters & HPC

Career Techniques • New York (NY)

Hybrid
USD 200,000 - 300,000
Senior AI GPU Cluster Architect
Senior AI GPU Cluster Architect

STN Inc • San Francisco (CA)

On-site
USD 180,000 - 240,000