Senior AI Infrastructure Engineer — GPU Clusters

Nvidia Corporation

Santa Clara (CA)

On-site

USD 152,000 - 287,500

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity
Benefits

Job summary

NVIDIA is hiring experienced software engineers to scale up its AI Infrastructure, focusing on cluster operations, operator development, and GPU resource scheduling.

You will help design and operate production systems for large-scale GPU clusters used by a variety of AI workloads, ensuring reliability and performance.

We value out-of-the-box problem solvers with strong execution bias and a solid foundation in Go or Python, data structures, and algorithms.

Qualifications

  • BS in CS, Eng, Physics or Math or equivalent experience.
  • 5+ years in software engineering on large-scale production systems.
  • Strong knowledge of Go or Python and data structures/algorithms.

Responsibilities

  • Be part of a DGX Cloud team responsible for production systems that enable large scalable GPU clusters.
  • Design and develop a massively distributed platform to identify, diagnose and remediate non-performant GPU assets.
  • Collaborate across NVIDIA teams to ensure production AI clusters run reliably with maximum performance.

Skills

Go
Python
Distributed systems
Algorithms
Communication

Education

BS in Computer Science, Engineering, Physics, Mathematics or equivalent

Tools

Kubernetes
Slurm
Base Command Manager

Job description

NVIDIA is hiring experienced software engineers to scale up its AI Infrastructure, focusing on cluster operations, operator development, and GPU resource scheduling.

You will help design and operate production systems for large-scale GPU clusters used by a variety of AI workloads, ensuring reliability and performance.

We value out-of-the-box problem solvers with strong execution bias and a solid foundation in Go or Python, data structures, and algorithms.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI Infra Engineer — Scalable GPU Clusters
Senior AI Infra Engineer — Scalable GPU Clusters

2100 NVIDIA USA • Santa Clara (CA)

On-site
USD 152,000 - 288,000
Equity
Benefits
Senior AI Infrastructure Engineer — Scale GPU Clusters Remote
Senior AI Infrastructure Engineer — Scale GPU Clusters Remote

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 280,000 - 420,000
Equity options
Health, vision, dental benefits
Unlimited PTO
+2
Lead AI Infrastructure Architect for Large-Scale GPU Clusters
Lead AI Infrastructure Architect for Large-Scale GPU Clusters

NVIDIA • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Equity and benefits
Senior Full-Stack Engineer — AI Infra & GPU Cloud (Equity)
Senior Full-Stack Engineer — AI Infra & GPU Cloud (Equity)

Nvidia Corporation • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Senior AI Infrastructure Architect — Enterprise GPU Clusters
Senior AI Infrastructure Architect — Enterprise GPU Clusters

NVIDIA • California (MO)

On-site
USD 184,000 - 288,000
Equity
Benefits
Senior AI Infrastructure Engineer | Scale GPU Clusters
Senior AI Infrastructure Engineer | Scale GPU Clusters

Fuel Talent LLC • Seattle (WA)

Hybrid
USD 126,000 - 189,000
Senior AI Infrastructure Architect – GPU Clusters
Senior AI Infrastructure Architect – GPU Clusters

NVIDIA • California (MO)

On-site
USD 184,000 - 357,000
Equity
Benefits
Senior AI Infrastructure Engineer — Scale GPU Clusters
Senior AI Infrastructure Engineer — Scale GPU Clusters

AI Breaking Wire • San Francisco (CA)

On-site
USD 280,000 - 400,000
Equity
Medical, dental, and vision benefits
Unlimited PTO
+2
Senior AI Systems Tools Engineer - GPU Clusters
Senior AI Systems Tools Engineer - GPU Clusters

NVIDIA Gruppe • Santa Clara (CA)

Hybrid
USD 184,000 - 357,000
Senior Cloud Infra Engineer: GPU Cluster Automation
Senior Cloud Infra Engineer: GPU Cluster Automation

NVIDIA • Santa Clara (CA)

On-site
USD 184,000 - 288,000
Equity
Benefits