Senior AI Infra Engineer: Kubernetes & Distributed Systems

NVIDIA

Santa Clara (CA)

Hybrid

USD 184,000 - 357,000

Full time

13 days ago
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

NVIDIA is seeking a Senior Systems Software Engineer to advance distributed AI infrastructure on Kubernetes, containers, and performance optimization at scale in production. You will own hard technical problems across hyperscale clusters and collaborate with diverse teams to drive innovations.

You will work with researchers and upstream communities to shape AI workloads, contribute to OSS projects, and optimize AI runtime performance while controlling costs.

Qualifications

  • Bachelor’s or Master’s degree in Engineering or equivalent experience, ideally in Electrical, Computer Engineering, or Computer Science
  • 8+ years of experience in computer architecture, networking, storage systems, and accelerator-based platforms
  • Expertise in Kubernetes and familiarity with the broader CNCF ecosystem
  • Deep experience with large-scale, parallel, distributed accelerator systems and performance optimization of AI workloads
  • Experience with performance modeling and benchmarking for large-scale systems
  • Proficiency in Golang and/or Python
  • Strong familiarity with the NVIDIA software stack across training and inference
  • Expertise with at least one major public cloud provider (e.g., AWS, Azure, GCP, OCI)

Responsibilities

  • Lead end-to-end performance and scalability analyses across the Kubernetes-based accelerated runtime stack
  • Design upstream architectural changes to Kubernetes control plane for hyperscale clusters
  • Improve container startup and latency for scalable AI workloads on Kubernetes
  • Contribute to open-source projects related to Kubernetes for AI workloads
  • Advance scalability of confidential containers (CoCo) on Kubernetes
  • Model full AI-factory deployments using DSX and-scale simulators
  • Collaborate with researchers, developers, customers, and upstream groups to automate at-scale workload tests
  • Document methods and present findings at industry events and within NVIDIA OSS communities

Skills

Distributed systems
Performance optimization
CI/CD

Education

Bachelor's degree in Engineering or equivalent
Master's degree in Engineering or equivalent

Tools

Kubernetes
Grove
gateway-api-inference-extension
DSX

Job description

NVIDIA is seeking a Senior Systems Software Engineer to advance distributed AI infrastructure on Kubernetes, containers, and performance optimization at scale in production. You will own hard technical problems across hyperscale clusters and collaborate with diverse teams to drive innovations.

You will work with researchers and upstream communities to shape AI workloads, contribute to OSS projects, and optimize AI runtime performance while controlling costs.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Systems Software Engineer: AI Infra & Kubernetes
Senior Systems Software Engineer: AI Infra & Kubernetes

NVIDIA • Seattle (WA)

Hybrid
USD 184,000 - 357,000
Equity
Health benefits
Flexible work arrangement
+1
Senior AI Infra Systems Engineer - Equity
Senior AI Infra Systems Engineer - Equity

NVIDIA AI • Santa Clara (CA)

On-site
USD 180,000 - 260,000
Equity
Senior AI Infra Architect: Kubernetes at Scale (Equity)
Senior AI Infra Architect: Kubernetes at Scale (Equity)

NVIDIA Corporation • Santa Clara (CA), Northern (KY)

Hybrid
USD 224,000 - 357,000
Equity
Benefits
Hybrid Remote Senior AI Infra Engineer – Kubernetes Scale
Hybrid Remote Senior AI Infra Engineer – Kubernetes Scale

NVIDIA Gruppe • Santa Clara (CA)

Hybrid
USD 184,000 - 288,000
Equity and benefits package
Hybrid work preference with remote options
Senior AI Infra Engineer-Distributed GPU Clusters (Equity)
Senior AI Infra Engineer-Distributed GPU Clusters (Equity)

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 184,000 - 356,500
Senior AI Infrastructure Engineer – Scale & Equity
Senior AI Infrastructure Engineer – Scale & Equity

Nvidia Corporation • Santa Clara (CA)

On-site
USD 184,000 - 356,500
Equity
Benefits
Senior Go Engineer, AI Storage & Distributed Systems
Senior Go Engineer, AI Storage & Distributed Systems

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Senior Cloud AI Infra Engineer: Go/C, Kubernetes, GPUs
Senior Cloud AI Infra Engineer: Go/C, Kubernetes, GPUs

Nvidia • Santa Clara (CA)

On-site
USD 240,000 - 340,000
Competitive salaries
Comprehensive benefits package
Equity
Senior AI Inference Architect — Scale & Disaggregated Serving
Senior AI Inference Architect — Scale & Disaggregated Serving

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Equity
Comprehensive benefits package
Senior Kubernetes & GPU AI Infra Engineer
Senior Kubernetes & GPU AI Infra Engineer

Nvidia Corporation • Durham (NC)

On-site
USD 248,000 - 397,000
Equity
Benefits