Senior SRE — GPU Cloud, Large-Scale Clusters & Kubernetes

NVIDIA Corporation

Santa Clara (CA)

On-site

USD 168,000 - 334,000

Full time

7 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

NVIDIA in Santa Clara, CA is seeking a Senior Site Reliability Engineer to significantly impact the success of external customers running NVIDIA solutions and internal clusters used for research, operations, and next-generation projects.

You’ll contribute to deployments, handle incidents, design features for the Base Command Manager, and validate Slurm and Kubernetes configurations to meet real-world customer needs.

Qualifications

  • Bachelor's Degree or equivalent experience in Computer Science or related field.
  • 8+ years of experience in site reliability engineering and/or software development roles.
  • Fluency in Python.
  • In-depth knowledge of Linux and networking.

Responsibilities

  • Contributing to deployments and daily operations of large scale next-generation GPU platforms
  • Handling incidents in GPU clusters, bridging the gap between cluster operations and development
  • Designing and implementing small features in the Base Command Manager product to become intimately familiar with the workings of the product
  • Validating complex cluster configurations including Slurm and Kubernetes orchestrators for performance, scalability and resilience, ensuring they meet real-world customer scenarios

Skills

Python
Linux networking

Education

Bachelor's degree in Computer Science or related field

Tools

Kubernetes
InfiniBand
Spectrum-X

Job description

NVIDIA in Santa Clara, CA is seeking a Senior Site Reliability Engineer to significantly impact the success of external customers running NVIDIA solutions and internal clusters used for research, operations, and next-generation projects.

You’ll contribute to deployments, handle incidents, design features for the Base Command Manager, and validate Slurm and Kubernetes configurations to meet real-world customer needs.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior SRE, BCM/DGX Cloud - Scale GPU Clusters
Senior SRE, BCM/DGX Cloud - Scale GPU Clusters

NVIDIA • Santa Clara (CA)

On-site
USD 168,000 - 334,000
Equity
Benefits
Senior SRE — Cloud Gaming Reliability & Automation
Senior SRE — Cloud Gaming Reliability & Automation

2100 NVIDIA USA • Santa Clara (CA)

On-site
USD 168,000 - 270,000
Equity
Benefits
Senior SRE — Cloud, Edge AI & CDN Automation
Senior SRE — Cloud, Edge AI & CDN Automation

Nvidia Corporation in • Santa Clara (CA)

On-site
USD 168,000 - 265,000
Equity
Comprehensive benefits
Career growth opportunities
Senior SRE — Cloud Gaming Reliability & Automation
Senior SRE — Cloud Gaming Reliability & Automation

NVIDIA • California (MO)

On-site
USD 168,000 - 270,000
Equity
Benefits
Distinguished Site Reliability Engineer - Cloud
Distinguished Site Reliability Engineer - Cloud

2100 NVIDIA USA • United States

On-site
USD 320,000 - 489,000
Equity
Benefits
Senior Cloud-Native Engineer — Kubernetes & Slurm for Multi-Tenant GPUs
Senior Cloud-Native Engineer — Kubernetes & Slurm for Multi-Tenant GPUs

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 184,000 - 288,000
Equity options
Comprehensive benefits
Senior SRE — Cloud Gaming Reliability & Automation
Senior SRE — Cloud Gaming Reliability & Automation

NVIDIA • Santa Clara (CA)

On-site
USD 168,000 - 270,000
Equity
Benefits
Senior SRE: 24/7 GPU & Kubernetes Reliability
Senior SRE: 24/7 GPU & Kubernetes Reliability

Nvidia Corporation in • Austin (TX)

On-site
USD 208,000 - 334,000
Equity
Benefits
Cloud SRE Architect — AI-Driven CI/CD & Scale
Cloud SRE Architect — AI-Driven CI/CD & Scale

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 272,000 - 431,000
Equity
Benefits
Senior Site Reliability Engineer - HPC
Senior Site Reliability Engineer - HPC

NVIDIA • Austin (TX)

On-site
USD 152,000 - 242,000
Equity
Comprehensive benefits package