Senior Cloud-Native AI Data Center Engineer Kubernetes/Slurm

NVIDIA

California (MO)

On-site

USD 184,000 - 357,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Equity
Benefits

Job summary

NVIDIA is seeking a Senior Software Engineer for the CSP Engagements team to advance the cloud-native stack for datacenter products like GB200/GB300. You will define customer workflows, prototype enhancements, and debug complex Kubernetes and Slurm issues in multi-rack, multi-tenant AI datacenters.

You will collaborate across CSP and internal platform teams to drive architecture, write RFCs, and deliver demonstrations at customer sites and conferences.

Qualifications

  • 10+ years of professional software development in distributed systems.
  • Strong source-level Kubernetes and Slurm expertise.
  • Experience with GPUs in containerized clusters.
  • Customer-facing engineering or solutions-architect background.
  • Familiarity with CI/CD, observability, and infrastructure-as-code.
  • Excellent communication and problem-solving abilities.

Responsibilities

  • Perform deep-dive debugging of multi-rack, multi-tenant clusters and scheduler behavior.
  • Prototype feature extensions for Kubernetes operators and Slurm plugins.
  • Lead architecture reviews and RFCs with CSP and internal teams.
  • Create reproducible testbeds (Helm/Ansible/Terraform) and automate validation.
  • Produce technical collateral and demo scripts for on-sites and conferences.
  • Collaborate with AE, FAE, and Solution Architect teams to deliver solutions.

Skills

Kubernetes internals
Slurm integration
CI/CD
Observability
Infrastructure as Code
Go/Rust/Python tooling
GPU integration
Customer communication

Education

BS or MS in Computer Engineering/CS

Tools

Helm
Terraform
Ansible
GitHub Actions
Tekton
Prometheus/OpenTelemetry
Slurm

Job description

NVIDIA is seeking a Senior Software Engineer for the CSP Engagements team to advance the cloud-native stack for datacenter products like GB200/GB300. You will define customer workflows, prototype enhancements, and debug complex Kubernetes and Slurm issues in multi-rack, multi-tenant AI datacenters.

You will collaborate across CSP and internal platform teams to drive architecture, write RFCs, and deliver demonstrations at customer sites and conferences.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior AI Datacenter Software Engineer (Kubernetes/Slurm)
Senior AI Datacenter Software Engineer (Kubernetes/Slurm)

BranchFactor • Austin (TX), Northern (KY)

Hybrid
USD 184,000 - 357,000
Equity
Benefits
Senior Cloud & HPC DevOps Architect
Senior Cloud & HPC DevOps Architect

Embedded Shishya • United States

Remote
USD 180,000 - 260,000
Senior Software Engineer, Cloud-Native Stack – CSP Engagements
Senior Software Engineer, Cloud-Native Stack – CSP Engagements

NVIDIA • California (MO)

On-site
USD 184,000 - 357,000
Equity
Benefits
Senior Cloud AI Infra Engineer: Go/C, Kubernetes, GPUs
Senior Cloud AI Infra Engineer: Go/C, Kubernetes, GPUs

Nvidia • Santa Clara (CA)

On-site
USD 240,000 - 340,000
Competitive salaries
Comprehensive benefits package
Equity
Senior Slurm Support Engineer for AI & HPC - Equity
Senior Slurm Support Engineer for AI & HPC - Equity

Nvidia Corporation in • Austin (TX)

On-site
USD 108,000 - 207,000
Senior Software Engineer
Senior Software Engineer

BranchFactor • Austin (TX), Northern (KY)

Hybrid
USD 184,000 - 357,000
Equity
Benefits
Remote Cluster Ops Engineer - Kubernetes, Slurm, NVIDIA
Remote Cluster Ops Engineer - Kubernetes, Slurm, NVIDIA

Comet Cloud • Boston (MA)

Hybrid
USD 140,000 - 200,000
Senior Slurm Support Engineer for HPC/AI Clusters
Senior Slurm Support Engineer for HPC/AI Clusters

NVIDIA • Austin (TX)

On-site
USD 108,000 - 173,000
Equity
Benefits
Senior Cloud Systems Engineer - Scientific Computing PaaS
Senior Cloud Systems Engineer - Scientific Computing PaaS

NVIDIA • Santa Clara (CA)

On-site
USD 184,000 - 288,000
Equity
Benefits
Remote Senior Cloud Data Platform Engineer
Remote Senior Cloud Data Platform Engineer

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 184,000 - 288,000
Equity
Benefits