Senior SRE, BCM/DGX Cloud - Scale GPU Clusters

NVIDIA

Santa Clara (CA)

On-site

USD 168,000 - 334,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity
Benefits

Job summary

NVIDIA is seeking a Sr Site Reliability Engineer in Santa Clara, CA to help operate large-scale GPU platforms and deploy robust infrastructure. You will contribute to deployments and daily operations, and bridge gaps between cluster operations and development.

Required: 8+ years in SRE or software roles, a CS degree or equivalent, and Python fluency. Linux networking is essential; Kubernetes and InfiniBand/Spectrum-X experience are a plus. Equity and benefits are offered.

Qualifications

  • Bachelor's Degree or equivalent in Computer Science or related field.
  • 8+ years of experience in site reliability engineering and/or software development roles.
  • Fluency in Python and strong Linux/networking knowledge.

Responsibilities

  • Contributing to deployments and daily operations of large scale next-generation GPU platforms.
  • Handling incidents in GPU clusters, bridging operations and development.
  • Designing and implementing small features in the Base Command Manager product.
  • Validating complex cluster configurations including Slurm and Kubernetes for performance and resilience.

Skills

Python
Linux
C++
Networking

Education

Bachelor's Degree or equivalent in Computer Science or related field

Tools

Kubernetes
InfiniBand
Spectrum-X

Job description

NVIDIA is seeking a Sr Site Reliability Engineer in Santa Clara, CA to help operate large-scale GPU platforms and deploy robust infrastructure. You will contribute to deployments and daily operations, and bridge gaps between cluster operations and development.

Required: 8+ years in SRE or software roles, a CS degree or equivalent, and Python fluency. Linux networking is essential; Kubernetes and InfiniBand/Spectrum-X experience are a plus. Equity and benefits are offered.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior SRE, BCM/DGX Cloud - Scale GPU Clusters (Equity)
Senior SRE, BCM/DGX Cloud - Scale GPU Clusters (Equity)

Nvidia Corporation in • Santa Clara (CA)

On-site
USD 168,000 - 334,000
Equity
Benefits
Senior SRE - Scalable Infra & Reliability (Equity)
Senior SRE - Scalable Infra & Reliability (Equity)

NVIDIA • Durham (NC)

On-site
USD 224,000 - 431,250
Equity
Benefits
SRE: Hardware Infrastructure for Reliable, AI‑Driven Ops
SRE: Hardware Infrastructure for Reliable, AI‑Driven Ops

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 184,000 - 287,500
Senior GPU Infra SRE — Remote, Low-Level Linux & Scale
Senior GPU Infra SRE — Remote, Low-Level Linux & Scale

Luma • Redwood City (CA)

Remote
USD 180,000 - 240,000
Senior SRE: FinOps-Driven Infra, GPU & Scale
Senior SRE: FinOps-Driven Infra, GPU & Scale

Level AI • Mountain View (CA)

On-site
USD 180,000 - 240,000
Senior SRE: 24/7 GPU & Kubernetes Reliability
Senior SRE: 24/7 GPU & Kubernetes Reliability

Nvidia Corporation in • Austin (TX)

On-site
USD 208,000 - 334,000
Equity
Benefits
Senior SRE, AI Infrastructure & GPU Fleet Reliability
Senior SRE, AI Infrastructure & GPU Fleet Reliability

Hamilton Barnes Associates Limited • San Francisco (CA)

On-site
USD 297,500 - 402,500
Huge stock options
Company bonus
Unlimited PTO
+1
Senior Site Reliability Engineer, BCM - DGX Cloud
Senior Site Reliability Engineer, BCM - DGX Cloud

Nvidia Corporation in • Santa Clara (CA)

On-site
USD 168,000 - 334,000
Equity
Benefits
Senior Site Reliability Engineer, BCM - DGX Cloud
Senior Site Reliability Engineer, BCM - DGX Cloud

NVIDIA • Santa Clara (CA)

On-site
USD 168,000 - 334,000
Equity
Benefits
Senior SRE Platform Engineer - Global GPU & AI Infra
Senior SRE Platform Engineer - Global GPU & AI Infra

Bitdeer Group • San Jose (CA)

On-site
USD 120,000 - 160,000