Senior AI Infrastructure Engineer — Scalable GPU Clusters

NVIDIA AI

Santa Clara (CA)

On-site

USD 152,000 - 288,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

NVIDIA in Santa Clara is seeking experienced software engineers to scale up its AI infrastructure, focusing on production systems, cluster operations, and GPU resource scheduling. You will help design a massively distributed platform and work with cross-functional teams to ensure reliable, high-performance AI clusters.

Applicants should have 5+ years on large-scale systems, a BS in CS/Engineering or equivalent, and Go or Python experience; Kubernetes, Slurm and related tooling are a plus.

Qualifications

  • BS in Computer Science, Engineering, Physics, Mathematics or comparable degree.
  • Experience with systems programming languages (Go, Python) and solid data structures/algorithms.
  • 5+ years in similar role working on large-scale production systems.
  • Direct experience in software engineering within a highly technical organization.

Responsibilities

  • Be part of a DGX Cloud team responsible for production systems enabling large scalable GPU clusters.
  • Design and develop a massively distributed scalable platform to identify, diagnose and remediate non-performant GPU assets.
  • Collaborate across NVIDIA to ensure production AI clusters run reliably with maximum performance and incident management.

Skills

Go
Python
Distributed systems

Education

Bachelor's degree in Computer Science, Engineering, Physics, Mathematics or a comparable Degree

Tools

Kubernetes
Slurm
Base Command Manager

Job description

NVIDIA in Santa Clara is seeking experienced software engineers to scale up its AI infrastructure, focusing on production systems, cluster operations, and GPU resource scheduling. You will help design a massively distributed platform and work with cross-functional teams to ensure reliable, high-performance AI clusters.

Applicants should have 5+ years on large-scale systems, a BS in CS/Engineering or equivalent, and Go or Python experience; Kubernetes, Slurm and related tooling are a plus.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI Infra Engineer — Scalable GPU Clusters
Senior AI Infra Engineer — Scalable GPU Clusters

2100 NVIDIA USA • Santa Clara (CA)

On-site
USD 152,000 - 288,000
Equity
Benefits
Senior AI Infrastructure Engineer — GPU Clusters
Senior AI Infrastructure Engineer — GPU Clusters

Nvidia Corporation • Santa Clara (CA)

On-site
USD 152,000 - 288,000
Equity
Benefits
Senior Full-Stack Engineer, AI Infra & GPU Clusters
Senior Full-Stack Engineer, AI Infra & GPU Clusters

NVIDIA • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Equity
Benefits
Senior Full-Stack Engineer, AI Infra for GPU Clusters
Senior Full-Stack Engineer, AI Infra for GPU Clusters

NVIDIA • California (MO)

On-site
USD 184,000 - 357,000
Equity
Benefits
Senior Full-Stack Engineer, AI Infra & GPU Clusters
Senior Full-Stack Engineer, AI Infra & GPU Clusters

NVIDIA • Town of Texas (WI)

On-site
USD 184,000 - 357,000
Equity
Benefits
Senior Full-Stack AI Infra Engineer
Senior Full-Stack AI Infra Engineer

NVIDIA AI • Santa Clara (CA)

On-site
USD 184,000 - 288,000
Equity
Benefits
Senior Full-Stack Engineer, AI Infra & GPU Clusters (Equity)
Senior Full-Stack Engineer, AI Infra & GPU Clusters (Equity)

NVIDIA • Raleigh (NC)

On-site
USD 224,000 - 357,000
Equity
Benefits
Senior AI Infra Engineer - Distributed GPU Systems (Equity)
Senior AI Infra Engineer - Distributed GPU Systems (Equity)

NVIDIA • California (MO)

On-site
USD 170,000 - 288,000
Equity
Benefits
Senior Full-Stack Engineer — AI Infra & GPU Cloud (Equity)
Senior Full-Stack Engineer — AI Infra & GPU Cloud (Equity)

Nvidia Corporation • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Senior Full-Stack Engineer — AI Infra for GPU Clusters
Senior Full-Stack Engineer — AI Infra for GPU Clusters

Socket.dev • Washington

Hybrid
USD 152,000 - 288,000
Equity
Benefits