Senior NCX Infra & Day-2 Operations Engineer

NVIDIA AI

Seattle (WA)

On-site

USD 184,000 - 357,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Equity
Benefits package

Job summary

NVIDIA is seeking an NCX Senior Engineer to partner with NVIDIA Cloud Partners for reliable, production-grade infrastructure in large-scale AI deployments. You will drive Day 2 operations, health monitoring, and fleet lifecycle practices across GPU clusters and cloud environments.

Ideal candidates have extensive Linux/distributed systems experience, deep Kubernetes knowledge, and a proven ability to translate reference architectures into repeatable operating models.

Qualifications

  • BS, MS, or PhD in Computer Science, Computer/Electrical Engineering, or related technical field, or equivalent experience.
  • 8+ years of infrastructure engineering, SRE, DevOps, cloud platform engineering, or similar.
  • Strong experience operating Linux-based distributed systems and cloud infrastructure in production.

Responsibilities

  • Lead NCP Day 2 operational readiness efforts with NVIDIA Cloud Partners to set up systems, procedures, and automation for managing NVIDIA accelerated infrastructure post-deployment.
  • Build continuous infrastructure validation across GPU, CPU, storage, and network health on large-scale AI clusters.
  • Establish observability and telemetry with dashboards, alerts, and signals across compute, GPU, InfiniBand/RoCE networking, storage, Kubernetes, and AI workloads.

Skills

Kubernetes
Observability
Python
Go
Shell scripting

Education

Bachelor's degree in Computer Science or related field

Tools

Prometheus
Grafana
OpenTelemetry

Job description

NVIDIA is seeking an NCX Senior Engineer to partner with NVIDIA Cloud Partners for reliable, production-grade infrastructure in large-scale AI deployments. You will drive Day 2 operations, health monitoring, and fleet lifecycle practices across GPU clusters and cloud environments.

Ideal candidates have extensive Linux/distributed systems experience, deep Kubernetes knowledge, and a proven ability to translate reference architectures into repeatable operating models.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Cloud Infra Engineer, Day 2 Ops for AI Clusters
Senior Cloud Infra Engineer, Day 2 Ops for AI Clusters

NVIDIA AI • Indiana (PA)

On-site
USD 180,000 - 240,000
Senior NCX Engineer: Day 2 Ops & Observability
Senior NCX Engineer: Day 2 Ops & Observability

NVIDIA • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Equity
Benefits
NCX Senior Engineer
NCX Senior Engineer

NVIDIA AI • Indiana (PA)

On-site
USD 180,000 - 240,000
Senior AI Infrastructure Engineer – Scale & Equity
Senior AI Infrastructure Engineer – Scale & Equity

Nvidia Corporation • Santa Clara (CA)

On-site
USD 184,000 - 356,500
Equity
Benefits
Senior AI Infrastructure Engineer — Equity Eligible
Senior AI Infrastructure Engineer — Equity Eligible

2100 NVIDIA USA • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Equity
Benefits package
Competitive salaries
NCX Senior Engineer
NCX Senior Engineer

NVIDIA • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Equity
Benefits
NCX Senior Engineer
NCX Senior Engineer

NVIDIA AI • Seattle (WA)

On-site
USD 184,000 - 357,000
Equity
Benefits package
Senior Systems Software Engineer: AI Infra & Kubernetes
Senior Systems Software Engineer: AI Infra & Kubernetes

NVIDIA • Seattle (WA)

Hybrid
USD 184,000 - 357,000
Equity
Health benefits
Flexible work arrangement
+1
Senior MLOps Engineer — AI Infra & Scale Leader
Senior MLOps Engineer — AI Infra & Scale Leader

NVIDIA • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Equity
Benefits package
Senior Software Engineer - DGX Cloud Production Automation
Senior Software Engineer - DGX Cloud Production Automation

NVIDIA Gruppe • California (MO)

On-site
USD 184,000 - 357,000
Equity
Benefits