Senior Systems Software Engineer, Accelerated Kubernetes Performance and Scale - DGX Cloud

NVIDIA Gruppe

Santa Clara (CA)

Hybrid

USD 184,000 - 287,500

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity and benefits package
Hybrid work preference with remote options

Job summary

NVIDIA Gruppe is seeking a Senior Systems Software Engineer for its DGX Cloud organization, tasked with shaping AI infrastructure for large-scale, cost-effective deployments. Your focus will be on performance and scalability of AI workloads on Kubernetes-based systems.

The role requires extensive experience in computer architecture and Kubernetes. A strong background in performance optimization and distributed systems will be essential. Join NVIDIA to drive innovation in AI technology.

Qualifications

  • 8+ years of experience in computer architecture and networking.
  • Deep experience with large-scale, parallel, distributed systems.
  • Expertise with at least one major public cloud provider.

Responsibilities

  • Lead performance analysis across Kubernetes-based accelerated runtime stack.
  • Design architectural changes to the Kubernetes control plane.
  • Collaborate with AI researchers and customers to enhance performance.

Skills

Kubernetes
Golang
Python
Computer Architecture
Networking
Storage Systems
Performance Optimization
Distributed Systems

Education

Bachelor's or Master's degree in Engineering or equivalent

Job description

Senior Systems Software Engineer – DGX Cloud

Joining NVIDIA’s DGX Cloud organization, you will help shape AI infrastructure for large‑scale, cost‑effective deployments. This role focuses on performance, scalability, and cost optimization of AI workloads on Kubernetes‑based runtimes.

Responsibilities
  • Lead end‑to‑end performance and scalability analysis across the Kubernetes‑based accelerated runtime stack (control and data planes), including components such as GPU Operator, Network Operator, node‑feature‑discovery, topograph, dra‑driver‑nvidia‑gpu, and nvsentinel, tracking issues from orchestration down to the metal.
  • Design and contribute upstream architectural changes to the Kubernetes control plane and related projects to enable reliable operation at hyperscale cluster sizes.
  • Improve container startup and cold‑start latency to enable smooth, low‑latency inference scaling on Kubernetes across thousands of GPU nodes.
  • Assess, improve, and contribute to open‑source projects that make Kubernetes an outstanding platform for AI workloads (e.g., Grove and gateway‑api‑inference‑extension).
  • Advance scalability and performance of confidential containers (CoCo) on Kubernetes so encrypted inference workloads meet stringent efficiency and latency requirements in production.
  • Use DSX and related large‑scale simulation infrastructure to model full AI‑factory deployments and validate scalability across thousands of simulated GPUs, catching failures that emerge only at scale before hardware arrives.
  • Collaborate with AI researchers, developers, customers, and upstream communities to design automated, at‑scale workload tests, build monitoring/analysis tooling, and integrate continuous performance and scale testing into modern CI/CD workflows.
  • Document methods and results clearly and present findings internally and at industry events (e.g., KubeCon, GTC), while engaging with upstream groups (Kubernetes SIG Scalability, CNCF, and NVIDIA OSS communities) to influence AI workload performance and scalability directions.
Qualifications
  • Bachelor’s or Master’s degree in Engineering or equivalent experience, ideally in Electrical, Computer Engineering, or Computer Science.
  • 8+ years of experience in computer architecture, networking, storage systems, and accelerator‑based platforms.
  • Expertise in Kubernetes and familiarity with the broader CNCF ecosystem.
  • Deep experience with large‑scale, parallel, distributed accelerator systems and performance optimization of AI workloads.
  • Experience with performance modeling and benchmarking for large‑scale systems.
  • Proficiency in Golang and/or Python.
  • Strong familiarity with the NVIDIA software stack across training and inference.
  • Expertise with at least one major public cloud provider (e.g., AWS, Azure, GCP, or OCI).
Benefits
  • Base salary range: 184,000USD – 287,500USD (Level4) or 224,000USD – 356,500USD (Level5).
  • Equity and benefits package.
  • Hybrid work preference with remote options.
Equal Opportunity

NVIDIA is committed to fostering an inclusive work environment and is an equal opportunity employer. We do not discriminate on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status, or any other characteristic protected by law.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Systems Software Engineer, Accelerated Kubernetes Performance and Scale - DGX Cloud
Senior Systems Software Engineer, Accelerated Kubernetes Performance and Scale - DGX Cloud

NVIDIA • Seattle (WA)

Hybrid
USD 184,000 - 357,000
Equity
Benefits
Systems Software Engineer, Kubernetes Scale - DGX Cloud
Systems Software Engineer, Kubernetes Scale - DGX Cloud

NVIDIA • Indiana (PA)

On-site
USD 100,000 - 135,000
Senior AI Infrastructure Engineer - DGX Cloud
Senior AI Infrastructure Engineer - DGX Cloud

NVIDIA • Santa Clara (CA)

On-site
USD 152,000 - 288,000
Equity
Benefits
Senior Software Engineer - DGX Cloud
Senior Software Engineer - DGX Cloud

NVIDIA • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Equity
Comprehensive benefits
Competitive salary
Senior Software Engineer - DGX Cloud
Senior Software Engineer - DGX Cloud

NVIDIA Gruppe • Seattle (WA)

On-site
USD 184,000 - 357,000
Equity
Benefits package
Performance bonuses
Senior Software Engineer - DGX Cloud
Senior Software Engineer - DGX Cloud

NVIDIA • Washington

On-site
USD 184,000 - 357,000
Equity
Competitive salary
Benefits package
Senior Software Engineer - DGX Cloud
Senior Software Engineer - DGX Cloud

Thomas To • Seattle (WA)

On-site
USD 184,000 - 357,000
Senior Performance Engineer - DGX Cloud
Senior Performance Engineer - DGX Cloud

NVIDIA AI • Eugene (OR)

On-site
USD 224,000 - 432,000
Equity
Benefits
Senior Performance Engineer - DGX Cloud
Senior Performance Engineer - DGX Cloud

NVIDIA • Oregon (WI)

On-site
USD 224,000 - 432,000
Equity
Benefits
Senior Performance Engineer - DGX Cloud
Senior Performance Engineer - DGX Cloud

NVIDIA • Washington

On-site
USD 224,000 - 432,000
Equity
Benefits