Senior Cloud Infra Engineer, Day 2 Ops for AI Clusters

NVIDIA AI

Indiana (PA)

On-site

USD 180,000 - 240,000

Full time

3 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

NVIDIA is hiring an NCX Senior Engineer to drive Day 2 operations for NVIDIA Cloud Partners and ensure reliable, large-scale accelerated infrastructure in production. You will work across compute, storage, networking, and AI workloads in partnership with partner teams.

You will develop observability, automation, and lifecycle management practices, translate reference architectures into production runbooks, and lead ongoing validation to sustain high availability and performance for critical

Qualifications

  • BS, MS, or Ph.D. in CS/Engineering or related field, or equivalent experience.
  • 8+ years of infrastructure engineering, SRE, DevOps or related large-scale production roles.
  • Strong Linux-based distributed systems and cloud infrastructure experience.
  • Deep Kubernetes and container lifecycle knowledge for large multi-node environments.
  • Strong production observability including metrics, logging, dashboards and SLAs.
  • Experience crafting automation for infra lifecycle, failure detection, remediation and upgrades.
  • Strong networking fundamentals across compute, network and storage layers.
  • Programming and automation with Python, Go or shell scripting.

Responsibilities

  • Lead NCP Day 2 operational readiness efforts with partner teams to set up systems and runbooks.
  • Build continuous infrastructure validation across GPU/CPU/storage/network health for large AI clusters.
  • Establish observability, telemetry, dashboards, alerts across compute, networking, storage and AI workloads.
  • Develop automated detection and remediation workflows to minimize disruption.
  • Refine fleet lifecycle administration including driver firmware and patch management.
  • Operationalize NVIDIA reference architectures into production practices and criteria.
  • Define health signals, SLOs and validation mechanisms for reliability and readiness.
  • Create reusable operational frameworks, guides, and reference implementations.

Skills

Linux
Kubernetes
Observability
Python
Go

Education

BS/MS/PhD in CS/Engineering or related field

Tools

Prometheus
Grafana
OpenTelemetry
Alertmanager

Job description

NVIDIA is hiring an NCX Senior Engineer to drive Day 2 operations for NVIDIA Cloud Partners and ensure reliable, large-scale accelerated infrastructure in production. You will work across compute, storage, networking, and AI workloads in partnership with partner teams.

You will develop observability, automation, and lifecycle management practices, translate reference architectures into production runbooks, and lead ongoing validation to sustain high availability and performance for critical

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior NCX Infra & Day-2 Operations Engineer
Senior NCX Infra & Day-2 Operations Engineer

NVIDIA AI • Seattle (WA)

On-site
USD 184,000 - 357,000
Equity
Benefits package
Senior NCX Engineer: Day 2 Ops & Observability
Senior NCX Engineer: Day 2 Ops & Observability

NVIDIA • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Equity
Benefits
Senior Cloud Day-2 Solutions Architect (Partner Ops)
Senior Cloud Day-2 Solutions Architect (Partner Ops)

NVIDIA • Santa Clara (CA)

On-site
USD 224,000 - 357,000
Senior Solutions Architect — Day-2 Cloud Ops for AI GPUs
Senior Solutions Architect — Day-2 Cloud Ops for AI GPUs

Nvidia Corporation • Santa Clara (CA)

On-site
USD 224,000 - 357,000
Equity
Comprehensive benefits
NCX Senior Engineer
NCX Senior Engineer

NVIDIA AI • Indiana (PA)

On-site
USD 180,000 - 240,000
Senior AI Infrastructure Engineer – Scale & Equity
Senior AI Infrastructure Engineer – Scale & Equity

Nvidia Corporation • Santa Clara (CA)

On-site
USD 184,000 - 356,500
Equity
Benefits
Senior Cloud Data Services Engineer for AI Infrastructure
Senior Cloud Data Services Engineer for AI Infrastructure

Visa Hunt • United States

On-site
USD 184,000 - 288,000
Senior Cloud Infrastructure Engineer – Platform Automation
Senior Cloud Infrastructure Engineer – Platform Automation

NVIDIA Corporation • United States

Remote
USD 168,000 - 322,000
Equity
Benefits
Senior AI Infrastructure Engineer — Equity Eligible
Senior AI Infrastructure Engineer — Equity Eligible

2100 NVIDIA USA • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Equity
Benefits package
Competitive salaries
Senior MLOps Engineer — AI Infra & Scale Leader
Senior MLOps Engineer — AI Infra & Scale Leader

NVIDIA • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Equity
Benefits package