Senior AI Infra & Kubernetes Architect

NVIDIA Corporation

Durham (NC)

On-site

USD 272,000 - 431,000

Full time

2 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Equity
Benefits

Job summary

NVIDIA Corporation in Durham, NC is seeking experienced software engineers with Kubernetes expertise to scale AI infrastructure. You will contribute to production GPU clusters, implement GPU resource scheduling on Kubernetes, and build monitoring and health management to maximize reliability and performance.

This role requires 15+ years in a similar environment, strong knowledge of Go or Python, solid data structures and algorithms, and proven impact in a highly technical setting.

Qualifications

  • Direct software engineering experience in a highly technical organization with demonstrable impact.
  • Software development experience with Kubernetes APIs and frameworks, not just operating a cluster.
  • 15+ years in similar role and experience on large-scale production systems.
  • BS in Computer Science, Engineering, Physics, Mathematics or comparable degree or equivalent experience.
  • Technical knowledge including a systems programming language (Go, Python) and solid understanding of data structures and algorithms.

Responsibilities

  • Be part of the DGX Cloud team responsible for production systems enabling large scalable GPU clusters for AI workloads.
  • Develop software related to scheduling GPU resources on Kubernetes.
  • Implement monitoring and health management to maximize reliability, availability, and performance of AI infrastructure.
  • Ingest and correlate data from GPU hardware diagnostics, cluster telemetry, and network data.
  • Collaborate with teams across NVIDIA to ensure production clusters run reliably and with maximum performance.
  • Evaluate system failures and drive improvements through a formal incident management process.

Skills

Kubernetes experience
Go
Python
Data structures & algorithms
Strong communication

Education

Bachelor's degree in CS or related field

Job description

NVIDIA Corporation in Durham, NC is seeking experienced software engineers with Kubernetes expertise to scale AI infrastructure. You will contribute to production GPU clusters, implement GPU resource scheduling on Kubernetes, and build monitoring and health management to maximize reliability and performance.

This role requires 15+ years in a similar environment, strong knowledge of Go or Python, solid data structures and algorithms, and proven impact in a highly technical setting.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI Infrastructure Engineer: Kubernetes & GPU Clusters
AI Infrastructure Engineer: Kubernetes & GPU Clusters

NVIDIA • United States

Remote
USD 184,000 - 288,000
Senior Cloud AI Infra Engineer: Go/C, Kubernetes, GPUs
Senior Cloud AI Infra Engineer: Go/C, Kubernetes, GPUs

Nvidia • Santa Clara (CA)

On-site
USD 240,000 - 340,000
Competitive salaries
Comprehensive benefits package
Equity
Senior AI Infra Architect: Kubernetes at Scale (Equity)
Senior AI Infra Architect: Kubernetes at Scale (Equity)

NVIDIA Corporation • Santa Clara (CA), Northern (KY)

Hybrid
USD 224,000 - 357,000
Equity
Benefits
AI Infra Engineer – SRE (Kubernetes)
AI Infra Engineer – SRE (Kubernetes)

Berrybytes • United States

On-site
USD 110,000 - 150,000
Senior Systems Software Engineer: AI Infra & Kubernetes
Senior Systems Software Engineer: AI Infra & Kubernetes

NVIDIA • Seattle (WA)

Hybrid
USD 184,000 - 357,000
Equity
Health benefits
Flexible work arrangement
+1
Senior AI Infrastructure Engineer - GPU & Kubernetes
Senior AI Infrastructure Engineer - GPU & Kubernetes

HCL Technologies Limited • California (MO)

On-site
USD 120,000 - 180,000
401(k) retirement plan
Paid time off (PTO)
Paid holidays
+1
Senior Software Engineer, Distributed Systems Engineer - DGX Cloud
Senior Software Engineer, Distributed Systems Engineer - DGX Cloud

NVIDIA • United States

Remote
USD 184,000 - 288,000
Hybrid Remote Senior AI Infra Engineer – Kubernetes Scale
Hybrid Remote Senior AI Infra Engineer – Kubernetes Scale

NVIDIA Gruppe • Santa Clara (CA)

Hybrid
USD 184,000 - 287,500
Equity and benefits package
Hybrid work preference with remote options
Senior Systems Software Engineer - GPU-Driven Kubernetes
Senior Systems Software Engineer - GPU-Driven Kubernetes

NVIDIA AI • Seattle (WA)

On-site
USD 180,000 - 260,000
Equity
Benefits
Principal Software Engineer, Distributed Systems Engineer - DGX Cloud
Principal Software Engineer, Distributed Systems Engineer - DGX Cloud

NVIDIA Corporation • Durham (NC)

On-site
USD 272,000 - 431,000
Equity
Benefits