Senior Software Engineer - Distributed AI Infra (Kubernetes)

Engg

Santa Clara (CA)

On-site

USD 184,000 - 357,000

Full time

2 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Equity
Benefits

Job summary

NVIDIA is hiring experienced software engineers with Kubernetes experience to help scale AI Infrastructure. You will contribute to production systems for large GPU clusters and develop scheduling software on Kubernetes, with a focus on reliability and high performance.

Ideal candidates have 8+ years in a similar role, strong Go or Python skills, and a track record in large-scale production systems. This position offers equity and benefits, with base pay varying by level and location.

Qualifications

  • Direct experience in a software engineering role within a highly technical organization with demonstrable impact.
  • Software development experience with kubernetes APIs and frameworks, not just cluster operation.
  • Strong communication skills and ability to coordinate across teams and geographies.
  • 8+ years in a similar role and experience on large-scale production systems.
  • Technical knowledge including Go or Python and solid understanding of data structures and algorithms.

Responsibilities

  • Be part of a DGX Cloud team responsible for production systems enabling large scalable GPU clusters for AI workloads.
  • Develop custom software related to scheduling GPU resources on Kubernetes.
  • Implement monitoring and health management to ensure reliability, availability, and scalability of GPU assets.
  • Process multiple data streams from hardware diagnostics to cluster and network telemetry.
  • Collaborate across NVIDIA teams to ensure production AI clusters run reliably with maximum performance.
  • Evaluate failures and improve services via a defined incident management process.

Skills

Kubernetes APIs
Go
Python
Distributed systems

Education

BS in Computer Science, Engineering, Physics, Mathematics or comparable degree

Tools

Slurm
Bright Cluster Manager
GPU scheduling tooling

Job description

NVIDIA is hiring experienced software engineers with Kubernetes experience to help scale AI Infrastructure. You will contribute to production systems for large GPU clusters and develop scheduling software on Kubernetes, with a focus on reliability and high performance.

Ideal candidates have 8+ years in a similar role, strong Go or Python skills, and a track record in large-scale production systems. This position offers equity and benefits, with base pay varying by level and location.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Principal Kubernetes & AI Infrastructure Engineer
Principal Kubernetes & AI Infrastructure Engineer

NVIDIA • Durham (NC)

On-site
USD 272,000 - 431,000
Equity
Benefits
Senior AI Infra & Kubernetes Architect
Senior AI Infra & Kubernetes Architect

NVIDIA Corporation • Durham (NC)

On-site
USD 272,000 - 431,000
Equity
Benefits
AI Infrastructure Architect: Kubernetes & GPU Scaling
AI Infrastructure Architect: Kubernetes & GPU Scaling

NVIDIA • United States

Remote
USD 272,000 - 431,000
Senior Software Engineer, Distributed Systems Engineer - DGX Cloud
Senior Software Engineer, Distributed Systems Engineer - DGX Cloud

NVIDIA • United States

Remote
USD 184,000 - 288,000
Principal Software Engineer, Distributed Systems Engineer - DGX Cloud
Principal Software Engineer, Distributed Systems Engineer - DGX Cloud

NVIDIA • Durham (NC)

On-site
USD 272,000 - 431,000
Equity
Benefits
Senior Systems Software Engineer - GPU-Driven Kubernetes
Senior Systems Software Engineer - GPU-Driven Kubernetes

NVIDIA AI • Seattle (WA)

On-site
USD 180,000 - 260,000
Equity
Benefits
Principal Software Engineer, Distributed Systems Engineer - DGX Cloud
Principal Software Engineer, Distributed Systems Engineer - DGX Cloud

NVIDIA Corporation • Durham (NC)

On-site
USD 272,000 - 431,000
Equity
Benefits
Senior AI Infra Engineer - Kubernetes Scale & Performance
Senior AI Infra Engineer - Kubernetes Scale & Performance

NVIDIA Corporation • Santa Clara (CA)

Hybrid
USD 184,000 - 357,000
Equity
Benefits
Hybrid work
AI Infrastructure Engineer: Kubernetes & GPU Clusters
AI Infrastructure Engineer: Kubernetes & GPU Clusters

NVIDIA • United States

Remote
USD 184,000 - 288,000
Principal Software Engineer, Distributed Systems Engineer - DGX Cloud
Principal Software Engineer, Distributed Systems Engineer - DGX Cloud

NVIDIA • United States

Remote
USD 272,000 - 431,000