Senior AI Infra Engineer — Scalable GPU Clusters

2100 NVIDIA USA

Santa Clara (CA)

On-site

USD 152,000 - 287,500

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity
Benefits

Job summary

NVIDIA in Santa Clara is hiring experienced software engineers to scale up its AI infrastructure, focusing on production systems for large GPU clusters and scalable platforms.

The role emphasizes distributed systems design, cluster management, and performance tuning, with a BS in CS/Engineering and 5+ years in similar roles. Base salary ranges are provided, equity eligible, and applications run through Aug 1, 2026.

Qualifications

  • Significant software engineering experience in a highly technical organization with demonstrable impact.
  • Strong communication skills and ability to coordinate across multi-functional teams and geographies.
  • 5+ years in a similar role and experience on large-scale production systems.
  • BS in Computer Science, Engineering, Physics, Mathematics or a comparable degree or equivalent experience.
  • Proficiency in a systems programming language (Go, Python) and solid understanding of data structures and algorithms.

Responsibilities

  • Be part of a DGX Cloud team responsible for production systems enabling large scalable GPU clusters for AI workloads.
  • Design and develop a massively distributed scalable platform to identify, diagnose and remediate non-performant GPU assets.
  • Collaborate with NVIDIA teams to ensure production AI clusters run reliably and with maximum performance; evaluate failures and improve services via incident management.

Skills

Distributed systems
Go/Python programming
Communication

Education

BS in Computer Science/Engineering/Physics/Mathematics

Tools

Kubernetes
Slurm

Job description

NVIDIA in Santa Clara is hiring experienced software engineers to scale up its AI infrastructure, focusing on production systems for large GPU clusters and scalable platforms.

The role emphasizes distributed systems design, cluster management, and performance tuning, with a BS in CS/Engineering and 5+ years in similar roles. Base salary ranges are provided, equity eligible, and applications run through Aug 1, 2026.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI Infra Engineer: GPU Clusters & Kubernetes
Senior AI Infra Engineer: GPU Clusters & Kubernetes

Intelliswift - An LTTS Company • Sunnyvale (CA)

On-site
USD 120,000 - 150,000
Senior AI Infra Engineer - Scalable Cloud Platform
Senior AI Infra Engineer - Scalable Cloud Platform

NVIDIA • Redmond (WA)

On-site
USD 184,000 - 357,000
Equity
Benefits
Senior AI Compute Engineer - HPC Infra & Linux
Senior AI Compute Engineer - HPC Infra & Linux

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 148,000 - 288,000
Senior AI Infrastructure Engineer — Scale GPU Clusters Remote
Senior AI Infrastructure Engineer — Scale GPU Clusters Remote

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 280,000 - 420,000
Equity options
Health, vision, dental benefits
Unlimited PTO
+2
Senior AI GPU Infra Engineer — Performance & Scale
Senior AI GPU Infra Engineer — Performance & Scale

Summit Group Solutions, LLC • United States

On-site
USD 150,000 - 350,000
Senior HPC-AI Systems Architect (Equity)
Senior HPC-AI Systems Architect (Equity)

Nvidia Corporation in • Santa Clara (CA)

On-site
USD 176,000 - 334,000
Senior AI Infrastructure Engineer for Scalable GPU Cloud
Senior AI Infrastructure Engineer for Scalable GPU Cloud

NVIDIA • Santa Clara (CA)

On-site
USD 152,000 - 288,000
Equity
Benefits
Lead AI Infrastructure Architect for Large-Scale GPU Clusters
Lead AI Infrastructure Architect for Large-Scale GPU Clusters

NVIDIA • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Equity and benefits
Senior Software Engineer, Distributed Systems Engineer - DGX Cloud
Senior Software Engineer, Distributed Systems Engineer - DGX Cloud

2100 NVIDIA USA • Santa Clara (CA)

On-site
USD 152,000 - 288,000
Equity
Benefits
Senior Kubernetes & GPU AI Infra Engineer
Senior Kubernetes & GPU AI Infra Engineer

Nvidia Corporation • Durham (NC)

On-site
USD 248,000 - 397,000
Equity
Benefits