Senior Distributed Systems Engineer for AI GPU Clusters

NVIDIA Corporation

United States

On-site

USD 120,000 - 160,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

A leading technology company is hiring a Senior Software Engineer for their DGX Cloud to develop and manage production systems for GPU clusters used in AI workloads. Candidates should have over 5 years of software engineering experience, expertise in Kubernetes, and proficiency in programming languages like Go and Python. This position offers a unique opportunity to tackle complex challenges in high-performance computing.

Qualifications

  • 5+ years of software engineering experience in a technical organization.
  • Experience with Kubernetes APIs and frameworks.
  • Proficiency in a systems programming language (Go, Python).

Responsibilities

  • Develop and manage production systems for scalable GPU clusters.
  • Implement monitoring and health management for GPU assets.
  • Collaborate with cross-functional teams to improve AI clusters.

Skills

Software engineering experience
Proficiency in Go
Proficiency in Python
Experience with Kubernetes

Education

BS in Computer Science or equivalent

Job description

A leading technology company is hiring a Senior Software Engineer for their DGX Cloud to develop and manage production systems for GPU clusters used in AI workloads. Candidates should have over 5 years of software engineering experience, expertise in Kubernetes, and proficiency in programming languages like Go and Python. This position offers a unique opportunity to tackle complex challenges in high-performance computing.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI Infrastructure Engineer - GPU Clusters
Senior AI Infrastructure Engineer - GPU Clusters

NVIDIA • Town of Texas (WI)

On-site
USD 184,000 - 288,000
Equity
Benefits
Senior Software Engineer
Senior Software Engineer

NVIDIA Corporation • United States

On-site
USD 120,000 - 160,000
Senior AI Infra Engineer: GPU Clusters & Kubernetes
Senior AI Infra Engineer: GPU Clusters & Kubernetes

Intelliswift - An LTTS Company • Sunnyvale (CA)

On-site
USD 120,000 - 150,000
Senior AI Infra Engineer - Distributed GPU Systems (Equity)
Senior AI Infra Engineer - Distributed GPU Systems (Equity)

NVIDIA • California (MO)

On-site
USD 170,000 - 288,000
Equity
Benefits
Senior Full-Stack Engineer — AI Infra for GPU Clusters
Senior Full-Stack Engineer — AI Infra for GPU Clusters

Socket.dev • Washington

Hybrid
USD 152,000 - 288,000
Equity
Benefits
Senior Software Engineer, DGX Cloud Kubernetes & AI Infra
Senior Software Engineer, DGX Cloud Kubernetes & AI Infra

Thomas To • Seattle (WA)

On-site
USD 184,000 - 357,000
Senior Cloud-Native Engineer, AI Datacenters & GPUs
Senior Cloud-Native Engineer, AI Datacenters & GPUs

NVIDIA • Redmond (WA)

On-site
USD 184,000 - 288,000
Senior Full-Stack AI Infra Engineer
Senior Full-Stack AI Infra Engineer

NVIDIA AI • Santa Clara (CA)

On-site
USD 184,000 - 288,000
Equity
Benefits
Senior DGX Cloud Production Engineer — AI Infra & SRE
Senior DGX Cloud Production Engineer — AI Infra & SRE

NVIDIA Gruppe • California (MO)

On-site
USD 168,000 - 334,000
Senior AI Infrastructure Engineer (DGX Cloud)
Senior AI Infrastructure Engineer (DGX Cloud)

NVIDIA • United States

On-site
USD 184,000 - 357,000