Senior GPU Platform SRE | Kubernetes, Slurm, BCM | Equity

Visa Hunt

United States

On-site

USD 168,000 - 334,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Equity
Benefits

Job summary

NVIDIA is seeking a Sr Site Reliability Engineer to help deploy and operate large-scale GPU platforms. You will bridge cluster operations with development, handle incidents, and design features in the Base Command Manager product. You will validate Slurm and Kubernetes configurations for performance, scale, and resilience.

Base salaries vary by level, with equity and benefits available. Applications accepted through August 27, 2026, for an existing vacancy.

Qualifications

  • Bachelor's Degree or equivalent experience in Computer Science or related field.
  • 8+ years of experience in site reliability engineering and/or software development roles.
  • Fluency in Python.
  • In-depth knowledge of Linux and networking.

Responsibilities

  • Contributing to deployments and daily operations of large scale next-generation GPU platforms.
  • Handling incidents in GPU clusters, bridging the gap between cluster operations and development.
  • Designing and implementing small features in the Base Command Manager product to understand the product.
  • Validating complex cluster configurations including Slurm and Kubernetes orchestrators for performance, scalability and resilience.

Skills

Python
Linux
Networking

Education

Bachelor's degree in Computer Science or related field

Tools

Kubernetes
InfiniBand
Spectrum-X

Job description

NVIDIA is seeking a Sr Site Reliability Engineer to help deploy and operate large-scale GPU platforms. You will bridge cluster operations with development, handle incidents, and design features in the Base Command Manager product. You will validate Slurm and Kubernetes configurations for performance, scale, and resilience.

Base salaries vary by level, with equity and benefits available. Applications accepted through August 27, 2026, for an existing vacancy.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer, BCM - DGX Cloud
Senior Site Reliability Engineer, BCM - DGX Cloud

Visa Hunt • United States

On-site
USD 168,000 - 334,000
Equity
Benefits
Senior Software Engineer, Cloud-Native Stack – CSP Engagements
Senior Software Engineer, Cloud-Native Stack – CSP Engagements

NVIDIA • California (MO)

On-site
USD 184,000 - 357,000
Equity
Benefits
Senior Site Reliability Engineer - HPC
Senior Site Reliability Engineer - HPC

NVIDIA Corporation • Durham (CA), Northern (KY)

On-site
USD 152,000 - 288,000
Senior Software Engineer, Cloud-Native Stack – CSP Engagements
Senior Software Engineer, Cloud-Native Stack – CSP Engagements

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Equity
Benefits
Senior Site Reliability Engineer - HPC
Senior Site Reliability Engineer - HPC

NVIDIA AI • Durham (NC)

On-site
USD 184,000 - 288,000
Equity
Benefits package
Principal Site Reliability Engineer
Principal Site Reliability Engineer

NVIDIA Gruppe • Santa Clara (CA)

Hybrid
USD 248,000 - 397,000
Equity
Benefits
Staff Site Reliability Engineer - AI Platform Runtime
Staff Site Reliability Engineer - AI Platform Runtime

Nvidia Corporation in • Santa Clara (CA)

Hybrid
USD 168,000 - 334,000
Staff Site Reliability Engineer - AI Platform Runtime
Staff Site Reliability Engineer - AI Platform Runtime

NVIDIA AI • Santa Clara (CA)

On-site
USD 168,000 - 334,000
Principal Site Reliability Engineer
Principal Site Reliability Engineer

NVIDIA • Santa Clara (CA)

Hybrid
USD 248,000 - 397,000
Senior GPU Infra SRE — Remote, Low-Level Linux & Scale
Senior GPU Infra SRE — Remote, Low-Level Linux & Scale

Luma • Redwood City (CA)

Remote
USD 180,000 - 240,000