Senior GPU Cloud Production Engineer — Kubernetes & Reliability

NVIDIA

Santa Clara (CA)

On-site

USD 272,000 - 431,250

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity
Benefits

Job summary

NVIDIA is seeking a Principal Software Engineer to define architecture and lead operations of large-scale GPU clusters. The role requires strong expertise in Kubernetes, Linux, and infrastructure automation.

Candidates should have over 15 years of experience with distributed systems, proven leadership in technical initiatives, and a degree in Computer Science. A competitive salary between $272,000 and $431,250 based on experience and location will be offered.

Qualifications

  • 15+ years of experience building and operating large-scale distributed systems or cloud infrastructure.
  • Deep experience with Kubernetes and production operations.
  • Strong programming experience in Go, Python, or similar.

Responsibilities

  • Define and execute the technical strategy for DGX Cloud cluster operations.
  • Lead design and implementation of systems for cluster lifecycle and readiness.
  • Identify and eliminate operational toil through software and automation.

Skills

Kubernetes
Linux
Infrastructure automation
Go
Python
Cloud infrastructure

Education

BS/MS in Computer Science or equivalent experience

Job description

NVIDIA is seeking a Principal Software Engineer to define architecture and lead operations of large-scale GPU clusters. The role requires strong expertise in Kubernetes, Linux, and infrastructure automation.

Candidates should have over 15 years of experience with distributed systems, proven leadership in technical initiatives, and a degree in Computer Science. A competitive salary between $272,000 and $431,250 based on experience and location will be offered.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Cloud Infra Engineer: GPU Cluster Automation
Senior Cloud Infra Engineer: GPU Cluster Automation

NVIDIA • Santa Clara (CA)

On-site
USD 184,000 - 288,000
Equity
Benefits
Senior Cloud Kubernetes Engineer for GPU AI Platform
Senior Cloud Kubernetes Engineer for GPU AI Platform

NVIDIA • Washington

On-site
USD 272,000 - 431,000
Senior Platform Engineer - Kubernetes & GPU AI Infra
Senior Platform Engineer - Kubernetes & GPU AI Infra

2100 NVIDIA USA • Seattle (WA)

On-site
USD 272,000 - 431,000
Equity
Benefits
Principal Software Engineer, DGX Cloud Production Engineering
Principal Software Engineer, DGX Cloud Production Engineering

Nvidia Corporation in • Santa Clara (CA)

On-site
USD 272,000 - 431,000
Equity
Senior Software Engineer, GPU Cloud Production & Automation
Senior Software Engineer, GPU Cloud Production & Automation

2100 NVIDIA USA • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Equity
Benefits
Principal Software Engineer, DGX Cloud Production Engineering
Principal Software Engineer, DGX Cloud Production Engineering

NVIDIA • Santa Clara (CA)

On-site
USD 272,000 - 431,250
Equity
Benefits
Senior Cloud-Native Engineer — Kubernetes & Slurm for Multi-Tenant GPUs
Senior Cloud-Native Engineer — Kubernetes & Slurm for Multi-Tenant GPUs

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 184,000 - 288,000
Equity options
Comprehensive benefits
Senior Software Engineer, DGX Cloud Production Engineering
Senior Software Engineer, DGX Cloud Production Engineering

Segment (Twilio) • Santa Clara (CA)

On-site
USD 184,000 - 288,000
Equity
Benefits
Kubernetes Platform Architect - Cloud & Fleet Ops (Equity)
Kubernetes Platform Architect - Cloud & Fleet Ops (Equity)

Nvidia Corporation in • Santa Clara (CA)

On-site
USD 272,000 - 431,000
Equity
Principal Cloud Engineer: Kubernetes & AI Infra, Equity
Principal Cloud Engineer: Kubernetes & AI Infra, Equity

NVIDIA Gruppe • Seattle (WA)

On-site
USD 272,000 - 431,000
Equity
Benefits package