Senior Full-Stack Engineer - AI Infra & GPU Clusters

NVIDIA

Washington (Washington County)

On-site

USD 184,000 - 357,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity and benefits

Job summary

NVIDIA is hiring experienced software engineers to scale up its AI infrastructure and production systems for large GPU clusters. You will contribute to diagnosing non-performant assets and improving system reliability across a broad stack including React, TypeScript, Golang, PostgreSQL, Temporal, Bazel and Kubernetes.

You will collaborate with multiple teams to ensure high performance, reliability and scalable deployment of AI workloads, leveraging cluster management and incident response

Qualifications

  • 5+ years in a similar role and experience on large-scale production systems.
  • BS in Computer Science or Engineering or equivalent experience.
  • Proficiency with React, TypeScript/JavaScript, and Golang.
  • Experience with distributed systems and production-grade software engineering.

Responsibilities

  • Be part of a DGX Cloud team for production systems enabling large scalable GPU clusters.
  • Design and develop a massively distributed platform to diagnose and remediate GPU assets.
  • Collaborate across teams to keep AI clusters reliable and high-performing.
  • Work across React, TypeScript, Golang, PostgreSQL, Temporal, Bazel, Kubernetes.

Skills

React
TypeScript/JavaScript
Golang
Distributed systems
Cluster operations

Education

BS in Computer Science or Engineering

Tools

Kubernetes
PostgreSQL
Temporal
Bazel
Slurm

Job description

NVIDIA is hiring experienced software engineers to scale up its AI infrastructure and production systems for large GPU clusters. You will contribute to diagnosing non-performant assets and improving system reliability across a broad stack including React, TypeScript, Golang, PostgreSQL, Temporal, Bazel and Kubernetes.

You will collaborate with multiple teams to ensure high performance, reliability and scalable deployment of AI workloads, leveraging cluster management and incident response

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Full-Stack Engineer, AI Infra & GPU Clusters
Senior Full-Stack Engineer, AI Infra & GPU Clusters

NVIDIA • Town of Texas (WI)

On-site
USD 184,000 - 357,000
Equity
Benefits
Senior Full-Stack Engineer, AI Infra & GPU Clusters (Equity)
Senior Full-Stack Engineer, AI Infra & GPU Clusters (Equity)

NVIDIA • Raleigh (NC)

On-site
USD 224,000 - 357,000
Equity
Benefits
Senior Full-Stack Engineer, AI Infra & GPU Clusters
Senior Full-Stack Engineer, AI Infra & GPU Clusters

NVIDIA • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Equity
Benefits
Senior Full-Stack Engineer, AI Infra for GPU Clusters
Senior Full-Stack Engineer, AI Infra for GPU Clusters

NVIDIA • California (MO)

On-site
USD 184,000 - 357,000
Equity
Benefits
Senior Full-Stack Engineer — AI Infra for GPU Clusters
Senior Full-Stack Engineer — AI Infra for GPU Clusters

Socket.dev • Washington

Hybrid
USD 152,000 - 288,000
Equity
Benefits
Senior Full-Stack AI Infra Engineer
Senior Full-Stack AI Infra Engineer

NVIDIA AI • Santa Clara (CA)

On-site
USD 184,000 - 288,000
Equity
Benefits
Senior Full-Stack Engineer — AI Infra & GPU Cloud (Equity)
Senior Full-Stack Engineer — AI Infra & GPU Cloud (Equity)

Nvidia Corporation • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Senior AI Infrastructure Engineer — GPU Clusters
Senior AI Infrastructure Engineer — GPU Clusters

Nvidia Corporation • Santa Clara (CA)

On-site
USD 152,000 - 288,000
Equity
Benefits
Senior AI Infra Engineer - Scalable GPU Clusters
Senior AI Infra Engineer - Scalable GPU Clusters

NVIDIA • Washington

On-site
USD 152,000 - 288,000
Equity
Benefits
Senior AI Infrastructure Engineer — Scalable GPU Clusters
Senior AI Infrastructure Engineer — Scalable GPU Clusters

NVIDIA AI • Santa Clara (CA)

On-site
USD 152,000 - 288,000