Senior AI Infra Engineer - Kubernetes Scale & Performance

NVIDIA Corporation

Santa Clara (CA)

Hybrid

USD 184,000 - 357,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Equity
Benefits
Hybrid work

Job summary

NVIDIA DGX Cloud seeks a Senior Systems Software Engineer with deep expertise in distributed systems, Kubernetes, containers, and performance and scalability. You will own hard technical problems at large scale, optimizing AI infrastructure in production and shaping how AI workloads run on Kubernetes across thousands of GPU nodes.

You will collaborate with researchers, developers, customers, and upstream communities, contribute to open-source projects, and design automated workload tests, while

Qualifications

  • Bachelor’s or Master’s degree in Engineering or equivalent experience.
  • 8+ years of experience in computer architecture, networking, storage systems, and accelerator-based platforms.
  • Expertise in Kubernetes and familiarity with the broader CNCF ecosystem.
  • Deep experience with large-scale, parallel, distributed accelerator systems and performance optimization of AI workloads.

Responsibilities

  • Lead end-to-end performance and scalability analysis across the Kubernetes-based accelerated runtime stack.
  • Design and contribute upstream architectural changes to the Kubernetes control plane.
  • Improve container startup and cold-start latency for thousands of GPU nodes.
  • Assess and contribute to open-source projects for AI workloads.
  • Advance scalability and performance of confidential containers on Kubernetes.

Skills

Kubernetes
Distributed systems
Golang/Python
Performance analysis

Education

Bachelor’s or Master’s degree in Engineering

Tools

Docker
Open-source ecosystems

Job description

NVIDIA DGX Cloud seeks a Senior Systems Software Engineer with deep expertise in distributed systems, Kubernetes, containers, and performance and scalability. You will own hard technical problems at large scale, optimizing AI infrastructure in production and shaping how AI workloads run on Kubernetes across thousands of GPU nodes.

You will collaborate with researchers, developers, customers, and upstream communities, contribute to open-source projects, and design automated workload tests, while

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Systems Software Engineer: AI Infra & Kubernetes
Senior Systems Software Engineer: AI Infra & Kubernetes

NVIDIA • Seattle (WA)

Hybrid
USD 184,000 - 357,000
Equity
Health benefits
Flexible work arrangement
+1
Hybrid Remote Senior AI Infra Engineer – Kubernetes Scale
Hybrid Remote Senior AI Infra Engineer – Kubernetes Scale

NVIDIA Gruppe • Santa Clara (CA)

Hybrid
USD 184,000 - 287,500
Equity and benefits package
Hybrid work preference with remote options
Senior Systems Engineer: Kubernetes & Scale for DGX Cloud
Senior Systems Engineer: Kubernetes & Scale for DGX Cloud

Engg • United States

Remote
USD 180,000 - 240,000
AI Infrastructure Architect: Kubernetes & GPU Scaling
AI Infrastructure Architect: Kubernetes & GPU Scaling

NVIDIA • United States

Remote
USD 272,000 - 431,000
Senior AI Infra & Kubernetes Architect
Senior AI Infra & Kubernetes Architect

NVIDIA Corporation • Durham (NC)

On-site
USD 272,000 - 431,000
Equity
Benefits
AI Infrastructure Engineer: Kubernetes & GPU Clusters
AI Infrastructure Engineer: Kubernetes & GPU Clusters

NVIDIA • United States

Remote
USD 184,000 - 288,000
Principal Kubernetes & AI Infrastructure Engineer
Principal Kubernetes & AI Infrastructure Engineer

NVIDIA • Durham (NC)

On-site
USD 272,000 - 431,000
Equity
Benefits
Senior Performance Engineer — AI Scale & Efficiency
Senior Performance Engineer — AI Scale & Efficiency

NVIDIA • Washington

On-site
USD 224,000 - 431,250
Equity
Benefits
Senior Performance Engineer for Scalable AI Workloads
Senior Performance Engineer for Scalable AI Workloads

NVIDIA • Oregon (WI)

On-site
USD 224,000 - 431,250
Equity
Benefits
Principal Cloud Engineer: Kubernetes & AI Infra, Equity
Principal Cloud Engineer: Kubernetes & AI Infra, Equity

NVIDIA Gruppe • Seattle (WA)

On-site
USD 272,000 - 431,000
Equity
Benefits package