Senior Platform & EngOps Engineer: GPU Cluster Automation

NVIDIA Corporation

Santa Clara (CA)

On-site

USD 176,000 - 276,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity options
Comprehensive benefits package

Job summary

NVIDIA Corporation in Santa Clara, CA is seeking a Senior Platform and EngOps Engineer for Cluster Operations. You will develop automated tools to manage GPU clusters and implement modern DevOps practices to ensure operational efficiency.

The ideal candidate has a strong background in computer engineering or a related field, with 8+ years of relevant experience and expertise in automation. Join us to contribute to groundbreaking advancements in AI and high-performance computing.

Qualifications

  • 8+ years of experience in deploying and administrating clusters and servers.
  • Experience with automation tools and scripting.
  • Proficient in managing high-performance applications.

Responsibilities

  • Develop automated tools for managing GPU clusters.
  • Implement DevOps tools for monitoring and maintenance.
  • Collaborate with Engineering teams for operations alignment.

Skills

Automation with Ansible
Python scripting
Shell scripting
Linux fundamentals
Understanding of operating systems
Knowledge of computer networks

Education

BS or MS in Computer Science, Computer Engineering, Electrical Engineering

Tools

DevOps tools
Slurm
GPU-focused hardware

Job description

NVIDIA Corporation in Santa Clara, CA is seeking a Senior Platform and EngOps Engineer for Cluster Operations. You will develop automated tools to manage GPU clusters and implement modern DevOps practices to ensure operational efficiency.

The ideal candidate has a strong background in computer engineering or a related field, with 8+ years of relevant experience and expertise in automation. Join us to contribute to groundbreaking advancements in AI and high-performance computing.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Platform & EngOps Engineer — GPU Clusters
Senior Platform & EngOps Engineer — GPU Clusters

NVIDIA AI • Santa Clara (CA)

On-site
USD 176,000 - 334,000
Equity
Benefits
Senior Platform and EngOps Engineer - Cluster Operations
Senior Platform and EngOps Engineer - Cluster Operations

NVIDIA AI • Santa Clara (CA)

On-site
USD 176,000 - 334,000
Equity
Benefits
Senior Platform and EngOps Engineer - Cluster Operations
Senior Platform and EngOps Engineer - Cluster Operations

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 176,000 - 276,000
Equity options
Comprehensive benefits package
Senior GPU HPC Cluster Engineer — Equity Eligible
Senior GPU HPC Cluster Engineer — Equity Eligible

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 152,000 - 242,000
Senior Cloud Infra Engineer: GPU Cluster Automation
Senior Cloud Infra Engineer: GPU Cluster Automation

NVIDIA • Santa Clara (CA)

On-site
USD 184,000 - 288,000
Equity
Benefits
Senior AI Infrastructure Engineer — Scalable GPU Clusters
Senior AI Infrastructure Engineer — Scalable GPU Clusters

NVIDIA AI • Santa Clara (CA)

On-site
USD 152,000 - 288,000
Senior AI Infrastructure Engineer — GPU Clusters
Senior AI Infrastructure Engineer — GPU Clusters

Nvidia Corporation • Santa Clara (CA)

On-site
USD 152,000 - 288,000
Equity
Benefits
Senior Full-Stack Engineer, AI Infra for GPU Clusters
Senior Full-Stack Engineer, AI Infra for GPU Clusters

NVIDIA • California (MO)

On-site
USD 184,000 - 357,000
Equity
Benefits
Senior HPC & GPU Cluster Architect — Scale & Automate
Senior HPC & GPU Cluster Architect — Scale & Automate

San Francisco Compute Company • San Francisco (CA)

Hybrid
USD 120,000 - 160,000
Generous equity grant
Competitive salary
Visa sponsorship
+6
Senior Software Engineer, GPU Cloud Production & Automation
Senior Software Engineer, GPU Cloud Production & Automation

2100 NVIDIA USA • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Equity
Benefits