Kubernetes Platform Lead — Production Reliability & Scale

NVIDIA

Washington (Washington County)

On-site

USD 176,000 - 333,500

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity
Benefits

Job summary

NVIDIA CFR is seeking a hands-on senior Kubernetes engineer to own the lifecycle and automation of the Kubernetes platform that supports GNI network systems. You will provide production support for platform-hosted services and partner with global teams on complex problems from design to production.

You will ensure cluster onboarding, upgrades, capacity planning, and recovery, while driving consistent engineering practices across US and Bangalore teams.

Qualifications

  • Bachelor's degree in Computer Science, Engineering, or related field or equivalent experience.
  • 8+ years building or operating production Kubernetes platforms, network infrastructure, or distributed systems.
  • Deep experience with Kubernetes at scale, including cluster lifecycle, upgrades, networking, storage, and recovery.
  • Proficiency in Go or Python.

Responsibilities

  • Design, build, and operate the Kubernetes platform powering GNI network automation, telemetry, and operations across data center, colocation, and cloud environments.
  • Own the lifecycle management for GNI Kubernetes environments, including cluster onboarding, upgrades, capacity, availability, and recovery.
  • Develop production-quality software and automation for cluster provisioning, validation, upgrades, remediation, and multi-cluster delivery through GitOps.
  • Provide production support for network services hosted on the platform, collaborating with Network Automation and service teams.
  • Diagnose complex Kubernetes platform and hosted-service failures and drive resolution from signal to verification.
  • Define production-readiness and observability standards, including health signals, capacity, alerts, runbooks, and recovery.
  • Participate in CFR’s production on-call rotation and lead incident response and corrective actions.

Skills

Kubernetes
Go
Python
GitOps
CI/CD
Automation
Production support

Education

Bachelor's degree in Computer Science or related field

Tools

Cluster API
Metal³
Git

Job description

NVIDIA CFR is seeking a hands-on senior Kubernetes engineer to own the lifecycle and automation of the Kubernetes platform that supports GNI network systems. You will provide production support for platform-hosted services and partner with global teams on complex problems from design to production.

You will ensure cluster onboarding, upgrades, capacity planning, and recovery, while driving consistent engineering practices across US and Bangalore teams.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Kubernetes Platform Engineer, Network Infrastructure
Senior Kubernetes Platform Engineer, Network Infrastructure

Thomas To • Santa Clara (CA)

On-site
USD 176,000 - 334,000
Equity and benefits
Senior Kubernetes Platform Architect: Global, Automated
Senior Kubernetes Platform Architect: Global, Automated

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 208,000 - 334,000
Equity
Benefits
Senior Kubernetes Platform Engineer - Global Network Infra
Senior Kubernetes Platform Engineer - Global Network Infra

NVIDIA • California (MO)

Hybrid
USD 176,000 - 334,000
Equity
Benefits
Senior Kubernetes Platform Engineer - Production Reliability
Senior Kubernetes Platform Engineer - Production Reliability

NVIDIA • Illinois

On-site
USD 208,000 - 334,000
Equity
Benefits
Senior Kubernetes Platform Engineer (Global Network)
Senior Kubernetes Platform Engineer (Global Network)

NVIDIA • Massachusetts

On-site
USD 176,000 - 334,000
Equity
Benefits
Senior Platform Engineer, Network Infrastructure - DGX Cloud
Senior Platform Engineer, Network Infrastructure - DGX Cloud

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 208,000 - 334,000
Equity
Benefits
Senior Platform Engineer, Network Infrastructure - DGX Cloud
Senior Platform Engineer, Network Infrastructure - DGX Cloud

NVIDIA • Washington

On-site
USD 176,000 - 334,000
Equity
Benefits
Senior Platform Engineer, Network Infrastructure - DGX Cloud
Senior Platform Engineer, Network Infrastructure - DGX Cloud

NVIDIA • Illinois

On-site
USD 208,000 - 334,000
Equity
Benefits
Senior Platform Engineer, Network Infrastructure - DGX Cloud
Senior Platform Engineer, Network Infrastructure - DGX Cloud

NVIDIA • California (MO)

Hybrid
USD 176,000 - 334,000
Equity
Benefits
Kubernetes Platform Architect - Cloud & Fleet Ops (Equity)
Kubernetes Platform Architect - Cloud & Fleet Ops (Equity)

Nvidia Corporation in • Santa Clara (CA)

On-site
USD 272,000 - 431,000
Equity