AI Infrastructure Engineer: Kubernetes & GPU Clusters

NVIDIA

United States

Remote

USD 184,000 - 288,000

Full time

4 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

NVIDIA in the United States is hiring experienced software engineers with Kubernetes experience to help scale AI infrastructure, focusing on GPU resource scheduling and cluster operations.

You will join the DGX Cloud team building production systems, implementing monitoring, health management, and reliable, high-performance AI clusters; you will collaborate across NVIDIA to resolve incidents and improve services.

Qualifications

  • Direct experience in a software engineering role within a highly technical organization with demonstrable impact.
  • Software development experience with Kubernetes APIs and frameworks, not just operating a cluster.
  • Strong communication, capable of cross-functional coordination across geographies.
  • 5+ years in similar role and experience on large-scale production systems.
  • Bachelor's degree in CS/Engineering/Physics/Math or equivalent.
  • Knowledge of a systems programming language (Go, Python) and solid data structures/algorithms.

Responsibilities

  • Work on production systems enabling scalable GPU clusters for AI workloads.
  • Develop scheduling software for GPU resources on Kubernetes.
  • Implement monitoring and health management for reliability and performance.
  • Collaborate across teams to ensure reliable AI clusters and incident response.

Skills

Kubernetes
Go
Python
Cluster ops
GPU scheduling
Incident mgmt

Education

BS in Computer Science, Engineering, Physics, Mathematics

Tools

Slurm
Bright Cluster Manager

Job description

NVIDIA in the United States is hiring experienced software engineers with Kubernetes experience to help scale AI infrastructure, focusing on GPU resource scheduling and cluster operations.

You will join the DGX Cloud team building production systems, implementing monitoring, health management, and reliable, high-performance AI clusters; you will collaborate across NVIDIA to resolve incidents and improve services.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior AI Infra & Kubernetes Architect
Senior AI Infra & Kubernetes Architect

NVIDIA Corporation • Durham (NC)

On-site
USD 272,000 - 431,000
Equity
Benefits
AI Infra Engineer: Kubernetes on Bare Metal GPUs (Remote)
AI Infra Engineer: Kubernetes on Bare Metal GPUs (Remote)

vCluster • New York (NY)

Hybrid
USD 150,000 - 200,000
Competitive Salary
Platinum-Level Insurance
Flexible Working Schedule
+1
Senior Systems Software Engineer: AI Infra & Kubernetes
Senior Systems Software Engineer: AI Infra & Kubernetes

NVIDIA • Seattle (WA)

Hybrid
USD 184,000 - 357,000
Equity
Health benefits
Flexible work arrangement
+1
Senior AI Infrastructure Engineer - GPU & Kubernetes
Senior AI Infrastructure Engineer - GPU & Kubernetes

HCL Technologies Limited • California (MO)

On-site
USD 120,000 - 180,000
401(k) retirement plan
Paid time off (PTO)
Paid holidays
+1
Hybrid Remote Senior AI Infra Engineer – Kubernetes Scale
Hybrid Remote Senior AI Infra Engineer – Kubernetes Scale

NVIDIA Gruppe • Santa Clara (CA)

Hybrid
USD 184,000 - 287,500
Equity and benefits package
Hybrid work preference with remote options
Principal Cloud Engineer: Kubernetes & AI Infra, Equity
Principal Cloud Engineer: Kubernetes & AI Infra, Equity

NVIDIA Gruppe • Seattle (WA)

On-site
USD 272,000 - 431,000
Equity
Benefits package
Senior Cloud AI Infra Engineer: Go/C, Kubernetes, GPUs
Senior Cloud AI Infra Engineer: Go/C, Kubernetes, GPUs

Nvidia • Santa Clara (CA)

On-site
USD 240,000 - 340,000
Competitive salaries
Comprehensive benefits package
Equity
Senior Software Engineer, Distributed Systems Engineer - DGX Cloud
Senior Software Engineer, Distributed Systems Engineer - DGX Cloud

NVIDIA • United States

Remote
USD 184,000 - 288,000
Senior Cloud Kubernetes Engineer for GPU AI Platform
Senior Cloud Kubernetes Engineer for GPU AI Platform

NVIDIA • Washington

On-site
USD 272,000 - 431,000
AI Infrastructure Engineer — GPU, Kubernetes & Automation
AI Infrastructure Engineer — GPU, Kubernetes & Automation

MARS-TECHNOMINDS, INC • Town of Florida (NY)

On-site
USD 100,000 - 130,000