Senior AI Infra Engineer - Distributed GPU Systems (Equity)

NVIDIA

California (MO)

On-site

USD 170,000 - 288,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity
Benefits

Job summary

NVIDIA is seeking experienced software engineers to scale AI infrastructure for DGX Cloud teams. You will work on production systems that manage large GPU clusters, focusing on performance, reliability, and incident management.

A BS in a technical field and 5+ years in a similar role are preferred, with strong Go/Python skills and distributed systems experience. You will collaborate across teams to improve platform reliability and performance, with equity and comprehensive benefits.

Qualifications

  • 5+ years in a software engineering role within a highly technical organization with demonstrable impact.
  • Strong communication skills and ability to coordinate across cross-functional teams and geographies.
  • Proficiency in Go and Python with solid data structures and algorithms.

Responsibilities

  • Maintain and scale production systems for large GPU clusters used in AI workloads.
  • Design and develop a massively distributed platform to identify, diagnose, and remediate non-performant GPU assets.
  • Collaborate with NVIDIA teams to ensure reliable and high-performance AI clusters and incident management.

Skills

Go
Python
Distributed systems
Communication skills

Education

BS in Computer Science, Engineering, Physics, Mathematics or comparable degree

Tools

Kubernetes
Slurm
Base Command Manager

Job description

NVIDIA is seeking experienced software engineers to scale AI infrastructure for DGX Cloud teams. You will work on production systems that manage large GPU clusters, focusing on performance, reliability, and incident management.

A BS in a technical field and 5+ years in a similar role are preferred, with strong Go/Python skills and distributed systems experience. You will collaborate across teams to improve platform reliability and performance, with equity and comprehensive benefits.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Full-Stack Engineer — AI Infra & GPU Cloud (Equity)
Senior Full-Stack Engineer — AI Infra & GPU Cloud (Equity)

Nvidia Corporation • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Senior Full-Stack AI Infra Engineer
Senior Full-Stack AI Infra Engineer

NVIDIA AI • Santa Clara (CA)

On-site
USD 184,000 - 288,000
Equity
Benefits
Senior AI Infrastructure Engineer for Scalable GPU Cloud
Senior AI Infrastructure Engineer for Scalable GPU Cloud

NVIDIA • Santa Clara (CA)

On-site
USD 152,000 - 288,000
Equity
Benefits
Senior Full-Stack Engineer, AI Infra & GPU Clusters (Equity)
Senior Full-Stack Engineer, AI Infra & GPU Clusters (Equity)

NVIDIA • Raleigh (NC)

On-site
USD 224,000 - 357,000
Equity
Benefits
Senior Full-Stack Engineer — AI Infra for GPU Clusters
Senior Full-Stack Engineer — AI Infra for GPU Clusters

Socket.dev • Washington

Hybrid
USD 152,000 - 288,000
Equity
Benefits
Senior Software Engineer
Senior Software Engineer

NVIDIA Corporation • United States

On-site
USD 120,000 - 160,000
Senior Software Engineer, Distributed Systems Engineer - DGX Cloud
Senior Software Engineer, Distributed Systems Engineer - DGX Cloud

2100 NVIDIA USA • Santa Clara (CA)

On-site
USD 152,000 - 288,000
Equity
Benefits
Senior Full-Stack Engineer, AI Infra for GPU Clusters
Senior Full-Stack Engineer, AI Infra for GPU Clusters

NVIDIA • California (MO)

On-site
USD 184,000 - 357,000
Equity
Benefits
Senior AI Infrastructure Engineer - DGX Cloud
Senior AI Infrastructure Engineer - DGX Cloud

NVIDIA • Santa Clara (CA)

On-site
USD 152,000 - 288,000
Equity
Benefits
Senior Software Engineer, Distributed Systems Engineer - DGX Cloud
Senior Software Engineer, Distributed Systems Engineer - DGX Cloud

NVIDIA • California (MO)

On-site
USD 170,000 - 288,000
Equity
Benefits