Senior AI Infrastructure Engineer - Kubernetes & Scale

NVIDIA

Seattle (WA)

Hybrid

USD 184,000 - 357,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity
Benefits

Job summary

NVIDIA is seeking a Senior Systems Software Engineer for the DGX Cloud org. You will tackle scaling AI infrastructure, building scalable Kubernetes-based runtimes, and optimizing performance across thousands of GPU nodes.

The role emphasizes collaboration with open-source communities and upstream projects. The team values innovation, cost efficiency, and driving impact at scale, with a preference for hybrid work and open to remote arrangements.

Qualifications

  • Bachelor’s or Master’s degree in Engineering or equivalent experience, ideally inElectrical, Computer Engineering, or Computer Science
  • 8+ years of experience in computer architecture, networking, storage systems and accelerator‑based platforms
  • Expertise in Kubernetes and familiarity with the broader CNCF ecosystem
  • Deep experience with large‑scale, parallel, distributed accelerator systems and performance optimization of AI workloads
  • Experience with performance modeling and benchmarking for large‑scale systems
  • Proficiency in Golang and/or Python
  • Strong familiarity with the NVIDIA software stack across training and inference
  • Expertise with at least one major public cloud provider (for example, AWS, Azure, GCP, or OCI)

Responsibilities

  • Lead end-to-end performance and scalability analysis across the Kubernetes‑based accelerated runtime stack.
  • Design and contribute upstream architectural changes to the Kubernetes control plane for hyperscale clusters.
  • Improve container startup and cold-start latency for scalable AI inference on Kubernetes across thousands of GPU nodes.
  • Contribute to open-source Kubernetes projects for AI workloads, focusing on scalability and multi-node training/inference.
  • Advance scalability and performance of confidential containers (CoCo) on Kubernetes.
  • Model AI-factory deployments using DSX to validate scalability on thousands of simulated GPUs.
  • Collaborate to design automated, at-scale workload tests and CI/CD performance testing.
  • Document results and present at industry events; engage with upstream groups to influence AI workload performance.

Skills

Kubernetes
Distributed systems
Golang
Python
Cloud platforms
Performance optimization
GPU/accelerator stacks

Education

Bachelor's or Master's in Engineering or CS

Job description

NVIDIA is seeking a Senior Systems Software Engineer for the DGX Cloud org. You will tackle scaling AI infrastructure, building scalable Kubernetes-based runtimes, and optimizing performance across thousands of GPU nodes.

The role emphasizes collaboration with open-source communities and upstream projects. The team values innovation, cost efficiency, and driving impact at scale, with a preference for hybrid work and open to remote arrangements.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Hybrid Remote Senior AI Infra Engineer – Kubernetes Scale
Hybrid Remote Senior AI Infra Engineer – Kubernetes Scale

NVIDIA Gruppe • Santa Clara (CA)

Hybrid
USD 184,000 - 288,000
Equity and benefits package
Hybrid work preference with remote options
Kubernetes Scale Engineer for AI Cloud Systems
Kubernetes Scale Engineer for AI Cloud Systems

NVIDIA • Indiana (PA)

On-site
USD 100,000 - 135,000
Senior Kubernetes AI Infra Engineer – DGX Cloud | Equity
Senior Kubernetes AI Infra Engineer – DGX Cloud | Equity

NVIDIA • Seattle (WA)

On-site
USD 184,000 - 357,000
Senior Software Engineer, DGX Cloud Kubernetes & AI Infra
Senior Software Engineer, DGX Cloud Kubernetes & AI Infra

Thomas To • Seattle (WA)

On-site
USD 184,000 - 357,000
Senior Cloud Kubernetes Engineer - GPU AI Infra
Senior Cloud Kubernetes Engineer - GPU AI Infra

NVIDIA AI • Seattle (WA)

On-site
USD 184,000 - 357,000
Equity
Benefits
Competitive salary
Senior Go Cloud Platform Engineer — Kubernetes & AI Infra
Senior Go Cloud Platform Engineer — Kubernetes & AI Infra

NVIDIA • Washington

On-site
USD 184,000 - 357,000
Equity
Competitive salary
Benefits package
Senior Software Engineer, GPU AI Infra on Kubernetes (Equity)
Senior Software Engineer, GPU AI Infra on Kubernetes (Equity)

NVIDIA Gruppe • Seattle (WA)

On-site
USD 184,000 - 357,000
Equity
Benefits package
Performance bonuses
Senior Kubernetes Systems Engineer - Scale & DGX Cloud
Senior Kubernetes Systems Engineer - Scale & DGX Cloud

Software Careers • Germany (OH)

On-site
USD 100,000 - 150,000
Senior Full-Stack Engineer — AI Infra & GPU Cloud (Equity)
Senior Full-Stack Engineer — AI Infra & GPU Cloud (Equity)

Nvidia Corporation • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Principal Cloud Engineer: Kubernetes & AI Infra, Equity
Principal Cloud Engineer: Kubernetes & AI Infra, Equity

NVIDIA Gruppe • Seattle (WA)

On-site
USD 272,000 - 431,000
Equity
Benefits package