Senior Platform Engineer, Network Infrastructure - DGX Cloud

NVIDIA Gruppe

Santa Clara (CA)

On-site

USD 208,000 - 333,500

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity
Benefits

Job summary

NVIDIA CFR is seeking a hands‑on senior Kubernetes platform engineer to own lifecycle, automation, and production support for the GNI Kubernetes fleet across US and Bangalore teams. You will design, build, and operate a scalable, observable platform with GitOps-driven delivery and strong on‑call discipline.

Join a critical backbone team in NVIDIA’s Global Network Infrastructure to drive reliability, scale, and efficiency across data centers and cloud environments.

Qualifications

  • Bachelor’s degree in Computer Science, Engineering, or related field, or equivalent experience.
  • 8+ years of experience building or operating production Kubernetes platforms, network infrastructure, or distributed systems.
  • Deep experience with Kubernetes at scale, including cluster lifecycle, upgrades, networking, storage, and recovery.
  • Proficiency in Go or Python.
  • Experience with GitOps, infrastructure as code, CI/CD, and automated production delivery.
  • Experience deploying and supporting network automation or telemetry services on Kubernetes.
  • Experience with production on‑call, incident response, root‑cause analysis, and driving corrective actions to completion.

Responsibilities

  • Design, build, and operate the Kubernetes platform powering GNI network automation, telemetry, and operations across data centers, colocation facilities, and cloud environments.
  • Own the lifecycle management for GNI Kubernetes environments, including cluster onboarding, upgrades, capacity, availability, and recovery.
  • Develop production‑quality software and automation for cluster provisioning, validation, upgrades, remediation, and multi‑cluster delivery through GitOps.
  • Provide production support for network services hosted on the platform, collaborating with Network Automation and service teams.
  • Diagnose complex Kubernetes platform and hosted‑service failures involving control‑plane health, cluster networking, storage, scheduling, and multi‑cluster dependencies; drive resolution.
  • Define production‑readiness and observability standards for the platform and hosted network services; establish health signals, capacity, alerts, runbooks, and recovery.
  • Participate in CFR’s production on‑call rotation, lead incident response and recovery, then drive corrective actions to completion.

Skills

Kubernetes at scale
Go or Python programming
GitOps practices
CI/CD
Incident response

Education

Bachelor's degree in Computer Science, Engineering, or related field

Tools

ClusterAPI (CAPI)
Metal3

Job description

Cloud Foundations Reliability (CFR) is part of NVIDIA’s Global Network Infrastructure (GNI) organization. We deploy, integrate, and operate the Kubernetes-based platform and shared services used to provision, monitor, and operate NVIDIA’s global network across data centers, colocation facilities, and cloud environments. The team owns the architecture and lifecycle of this platform, including cluster provisioning and upgrades, GitOps delivery, observability, capacity, and service enablement. We build software and automation to standardize how network platforms and services are deployed, scaled, and managed across environments.

We are looking for a hands‑on senior engineer to own the lifecycle and automation of the Kubernetes platform supporting GNI network systems. You will also provide production support for network services running on the platform, partnering with their engineering owners when issues or changes cross the platform boundary. You will take complex problems from design through production and remain accountable for the outcome. You will bring deep Kubernetes expertise and help establish consistent engineering practices across the US and Bangalore teams. This is a senior individual‑contributor role with end‑to‑end ownership and production responsibility.

What You’ll Be Doing:
  • Design, build, and operate the Kubernetes platform that powers GNI network automation, telemetry, and operations across data center, colocation, and cloud environments.
  • Own the lifecycle management for GNI Kubernetes environments, including cluster onboarding, upgrades, capacity, availability, and recovery.
  • Develop production‑quality software and automation for cluster provisioning, validation, upgrades, remediation, and safe multi‑cluster delivery through GitOps.
  • Provide production support for network services hosted on the platform, working with Network Automation and service teams that retain ownership of application architecture, code, and features.
  • Diagnose complex Kubernetes platform and hosted‑service failures involving control‑plane health, cluster networking, storage, scheduling, workload placement, and multi‑cluster dependencies. Drive issues from initial signal through verified resolution.
  • Define production‑readiness and observability standards for the platform and hosted network services, including health signals, capacity, alerts, runbooks, and recovery.
  • Participate in CFR’s production on‑call rotation, including scheduled after‑hours and weekend coverage. Lead incident response and recovery, then drive corrective actions to completion.
What We Need to See:
  • Bachelor’s degree in Computer Science, Engineering, or a related field, or equivalent experience.
  • 8+ years of experience building or operating production Kubernetes platforms, network infrastructure, or distributed systems.
  • Deep experience with Kubernetes at scale, including cluster lifecycle, upgrades, networking, storage, and recovery.
  • Proficiency in at least one general‑purpose programming language, such as Go or Python.
  • Experience with GitOps, infrastructure as code, CI/CD, and automated production delivery.
  • Experience deploying and supporting network automation or telemetry services on Kubernetes.
  • Experience with production on‑call, incident response, root‑cause analysis, and driving corrective actions to completion.
Ways to Stand Out From the Crowd:
  • Strong knowledge of IP routing, data center fabrics, and cloud networking is a great plus.
  • Experience designing and operating large, multi‑region Kubernetes fleets, including fleet‑wide upgrades and recovery.
  • Hands‑on experience with ClusterAPI (CAPI) and Metal3 for bare‑metal provisioning, cluster lifecycle, machine remediation, and upgrades.
  • Experience building Kubernetes controllers or operators in Go using custom resources and reconciliation patterns. Experience designing or operating network automation and telemetry services on Kubernetes at global scale.
  • Contributions to ClusterAPI, Metal3, or other open‑source Kubernetes infrastructure projects.

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 176,000USD–276,000USD for Level4, and 208,000USD–333,500USD for Level5.

You will also be eligible for equity and benefits.

Applications for this job will be accepted at least until July20,2026.

This posting is for an existing vacancy.

NVIDIA uses AI tools in its recruiting processes.

NVIDIA is committed to fostering an inclusive work environment and proud to be an equal‑opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Platform Engineer, Network Infrastructure - DGX Cloud
Senior Platform Engineer, Network Infrastructure - DGX Cloud

NVIDIA • Washington

On-site
USD 176,000 - 334,000
Equity
Benefits
Senior Platform Engineer, Network Infrastructure - DGX Cloud
Senior Platform Engineer, Network Infrastructure - DGX Cloud

NVIDIA • Illinois

On-site
USD 208,000 - 334,000
Equity
Benefits
Senior Platform Engineer, Network Infrastructure - DGX Cloud
Senior Platform Engineer, Network Infrastructure - DGX Cloud

NVIDIA • California (MO)

Hybrid
USD 176,000 - 334,000
Equity
Benefits
Senior Platform Engineer, Network Infrastructure - DGX Cloud
Senior Platform Engineer, Network Infrastructure - DGX Cloud

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 176,000 - 333,500
Equity
Benefits
Senior Platform Engineer, Network Infrastructure
Senior Platform Engineer, Network Infrastructure

NVIDIA AI • Indiana (PA)

On-site
USD 150,000 - 210,000
Senior Software Engineer, Networking DGX Cloud
Senior Software Engineer, Networking DGX Cloud

NVIDIA AI • New York (NY)

On-site
USD 200,000 - 391,000
Equity
Benefits
Senior Systems Software Engineer, Developer Productivity and Cloud Automation - GeForce NOW
Senior Systems Software Engineer, Developer Productivity and Cloud Automation - GeForce NOW

NVIDIA • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Equity
Comprehensive benefits
Senior Software Engineer, Networking DGX Cloud
Senior Software Engineer, Networking DGX Cloud

NVIDIA Gruppe • United States

On-site
USD 200,000 - 391,000
Equity and benefits
Senior Software Engineer - DGX Cloud
Senior Software Engineer - DGX Cloud

NVIDIA Gruppe • Seattle (WA)

On-site
USD 184,000 - 357,000
Equity
Benefits package
Performance bonuses
Senior Software Engineer - DGX Cloud
Senior Software Engineer - DGX Cloud

NVIDIA • Washington

On-site
USD 184,000 - 357,000
Equity
Competitive salary
Benefits package