Senior Platform & EngOps Engineer — GPU Clusters

NVIDIA AI

Santa Clara (CA)

On-site

USD 176,000 - 333,500

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity
Benefits

Job summary

NVIDIA is seeking EngOps and Platform Engineers to drive automation for GPU clusters interconnected via NVLink and InfiniBand. You will build deployment and monitoring tools and own cluster reliability across time zones.

The role demands 8+ years deploying clusters, strong Linux, Ansible, Python, and shell scripting, with a focus on collaboration with developers and testers. This is a full-time, on-site role in Santa Clara, CA, with equity and benefits.

Qualifications

  • BS or MS in Computer Science, Computer Engineering, Electrical Engineering, or related field, or equivalent experience.
  • 8+ years deploying and administering clusters, servers, switches, and related infrastructure.
  • Automation expertise with Ansible, Python and Shell Scripting.
  • Deep understanding of operating systems, networks, and high-performance apps.
  • Proven ability to work with developers and test engineers across teams and time zones.
  • Proficient with Linux fundamentals.

Responsibilities

  • Develop automated tools to deploy, provision, and maintain extensive GPU clusters.
  • Implement modern DevOps tools to automate updates, maintenance tasks, and monitor cluster availability.
  • Own daily cluster failures and issues, troubleshooting to maintain performance.
  • Manage rollout and rollback of software and firmware updates for clusters.
  • Collaborate with Engineering and Product Teams across time zones to align operations with project requirements.

Skills

Ansible
Python
Shell scripting
Linux
Cluster administration

Education

BS or MS in Computer Science / Computer Engineering / Electrical Engineering

Tools

Slurm

Job description

NVIDIA is seeking EngOps and Platform Engineers to drive automation for GPU clusters interconnected via NVLink and InfiniBand. You will build deployment and monitoring tools and own cluster reliability across time zones.

The role demands 8+ years deploying clusters, strong Linux, Ansible, Python, and shell scripting, with a focus on collaboration with developers and testers. This is a full-time, on-site role in Santa Clara, CA, with equity and benefits.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Platform & EngOps Engineer: GPU Cluster Automation
Senior Platform & EngOps Engineer: GPU Cluster Automation

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 176,000 - 276,000
Equity options
Comprehensive benefits package
Senior Platform and EngOps Engineer - Cluster Operations
Senior Platform and EngOps Engineer - Cluster Operations

NVIDIA AI • Santa Clara (CA)

On-site
USD 176,000 - 334,000
Equity
Benefits
Senior Platform and EngOps Engineer - Cluster Operations
Senior Platform and EngOps Engineer - Cluster Operations

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 176,000 - 276,000
Equity options
Comprehensive benefits package
Senior GPU HPC Cluster Engineer — Equity Eligible
Senior GPU HPC Cluster Engineer — Equity Eligible

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 152,000 - 242,000
Senior Full-Stack Engineer, AI Infra for GPU Clusters
Senior Full-Stack Engineer, AI Infra for GPU Clusters

NVIDIA • California (MO)

On-site
USD 184,000 - 357,000
Equity
Benefits
Senior GPU Infra Engineer: AI Clusters & OpenStack Lead
Senior GPU Infra Engineer: AI Clusters & OpenStack Lead

Hamilton Barnes Associates Limited • Town of Texas (WI)

On-site
USD 120,000 - 160,000
Potential equity/bonus
Senior AI Infrastructure Engineer — Scalable GPU Clusters
Senior AI Infrastructure Engineer — Scalable GPU Clusters

NVIDIA AI • Santa Clara (CA)

On-site
USD 152,000 - 288,000
Senior Platform Engineer – AI/ML Infra (Equity)
Senior Platform Engineer – AI/ML Infra (Equity)

NVIDIA Corporation • Santa Clara (CA), Northern (KY)

Hybrid
USD 200,000 - 322,000
Senior HPC AI Cluster Architect — Equity Eligible
Senior HPC AI Cluster Architect — Equity Eligible

NVIDIA Corporation • Santa Clara (CA), Northern (KY)

Hybrid
USD 176,000 - 334,000
Senior AI Infrastructure Engineer — GPU Clusters
Senior AI Infrastructure Engineer — GPU Clusters

Nvidia Corporation • Santa Clara (CA)

On-site
USD 152,000 - 288,000
Equity
Benefits