Senior Kubernetes Platform Engineer, Network Infrastructure

Thomas To

Santa Clara (CA)

On-site

USD 176,000 - 334,000

Full time

23 hours ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity and benefits

Job summary

NVIDIA’s Cloud Foundations Reliability (CFR) group seeks a senior Kubernetes Platform Engineer to own lifecycle management, automation, and production support for the GNI Kubernetes platform across US and Bangalore. You’ll drive design, provisioning, upgrades, and multi-cluster delivery with GitOps, ensuring reliability and observability.

You will collaborate with Network Automation and service teams, lead incident response, and enforce production-readiness standards.

Qualifications

  • Bachelor's degree in Computer Science, Engineering, or related field, or equivalent experience.
  • 8+ years of experience building or operating production Kubernetes platforms, network infrastructure, or distributed systems.
  • Deep experience with Kubernetes at scale, including cluster lifecycle, upgrades, networking, storage, and recovery.
  • Proficiency in at least one general-purpose programming language, such as Go or Python.
  • Experience with GitOps, infrastructure as code, CI/CD, and automated production delivery.
  • Experience deploying and supporting network automation or telemetry services on Kubernetes.
  • Experience with production on-call, incident response, root-cause analysis, and driving corrective actions to completion.

Responsibilities

  • Design, build, and operate the Kubernetes platform that powers GNI network automation, telemetry, and operations across data center, colocation, and cloud environments.
  • Own the lifecycle management for GNI Kubernetes environments, including cluster onboarding, upgrades, capacity, availability, and recovery.
  • Develop production-quality software and automation for cluster provisioning, validation, upgrades, remediation, and safe multi-cluster delivery through GitOps.
  • Provide production support for network services hosted on the platform, working with Network Automation and service teams that retain ownership of application architecture, code, and features.
  • Diagnose complex Kubernetes platform and hosted-service failures involving control-plane health, cluster networking, storage, scheduling, workload placement, and multi-cluster dependencies. Drive issues from initial signal through verified resolution.
  • Define production-readiness and observability standards for the platform and hosted network services, including health signals, capacity, alerts, runbooks, and recovery.
  • Participate in CFR's production on-call rotation, including scheduled after-hours and weekend coverage. Lead incident response and recovery, then drive corrective actions to completion.

Skills

Kubernetes
Go/Python
GitOps
Incident response

Education

Bachelor's degree in CS or related field

Tools

Cluster API (CAPI)
Metal3
CI/CD tooling

Job description

NVIDIA’s Cloud Foundations Reliability (CFR) group seeks a senior Kubernetes Platform Engineer to own lifecycle management, automation, and production support for the GNI Kubernetes platform across US and Bangalore. You’ll drive design, provisioning, upgrades, and multi-cluster delivery with GitOps, ensuring reliability and observability.

You will collaborate with Network Automation and service teams, lead incident response, and enforce production-readiness standards.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Kubernetes Platform Engineer - Global Network Infra
Senior Kubernetes Platform Engineer - Global Network Infra

NVIDIA • California (MO)

Hybrid
USD 176,000 - 334,000
Equity
Benefits
Senior Kubernetes Platform Architect: Global, Automated
Senior Kubernetes Platform Architect: Global, Automated

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 208,000 - 334,000
Equity
Benefits
Kubernetes Platform Lead — Production Reliability & Scale
Kubernetes Platform Lead — Production Reliability & Scale

NVIDIA • Washington

On-site
USD 176,000 - 334,000
Equity
Benefits
Senior Kubernetes Platform Engineer (Global Network)
Senior Kubernetes Platform Engineer (Global Network)

NVIDIA • Massachusetts

On-site
USD 176,000 - 334,000
Equity
Benefits
Senior Kubernetes Platform Engineer - Production Reliability
Senior Kubernetes Platform Engineer - Production Reliability

NVIDIA • Illinois

On-site
USD 208,000 - 334,000
Equity
Benefits
Senior Platform Engineer, Network Infrastructure - DGX Cloud
Senior Platform Engineer, Network Infrastructure - DGX Cloud

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 208,000 - 334,000
Equity
Benefits
Senior Platform Engineer, Network Infrastructure - DGX Cloud
Senior Platform Engineer, Network Infrastructure - DGX Cloud

NVIDIA • Illinois

On-site
USD 208,000 - 334,000
Equity
Benefits
Senior Platform Engineer, Network Infrastructure - DGX Cloud
Senior Platform Engineer, Network Infrastructure - DGX Cloud

NVIDIA • Washington

On-site
USD 176,000 - 334,000
Equity
Benefits
Senior Platform Engineer, Network Infrastructure - DGX Cloud
Senior Platform Engineer, Network Infrastructure - DGX Cloud

NVIDIA • California (MO)

Hybrid
USD 176,000 - 334,000
Equity
Benefits
Senior Platform Engineer, Network Infrastructure - DGX Cloud
Senior Platform Engineer, Network Infrastructure - DGX Cloud

Thomas To • Santa Clara (CA)

On-site
USD 176,000 - 334,000
Equity and benefits