Senior HPC-AI Cluster Architect (Equity)

NVIDIA

California (MO)

On-site

USD 176,000 - 334,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

NVIDIA is seeking an experienced HPC-AI Engineer to join the Networking Clusters Solutions Infrastructure team. You will design and run large-scale HPC/AI clusters, manage job workloads with Slurm and K8s, and automate deployment via CI/CD pipelines.

You’ll work with researchers and customers to craft workflows and advanced solutions, across Linux/Windows, networking, storage and virtualization layers. Base salary varies by level (approximately 176k–276k USD for Level 4 and 208k–333.5k USD for

Qualifications

  • Degree in CS or Engineering or equivalent experience; 8+ years in HPC/AI compute.
  • Knowledge of HPC/AI tech from CPU/GPU to high-speed interconnects.
  • Experience with job scheduling/workloads and orchestration tools like Slurm and K8s.
  • Excellent Windows/Linux networking knowledge and OS security basics.
  • Experience with Lustre/GPFS/Weka.io storage technologies.
  • Python and Bash scripting experience.
  • Comfort with Jenkins/Ansible/Puppet/Chef for automation.
  • Deep knowledge of InfiniBand and Ethernet networking protocols.
  • Experience with virtualization (VMware/Hyper-V/KVM/Citrix).
  • Familiarity with cloud platforms (AWS/Azure/Google Cloud).

Responsibilities

  • Design, implement and maintain large scale HPC/AI clusters with monitoring and alerting.
  • Manage Linux job/workload schedules and orchestration tools.
  • Develop and maintain CI/CD pipelines.
  • Create tooling to automate deployment and management of large-scale infra.
  • Deploy monitoring for servers, network and storage.
  • Troubleshoot from bare metal to application level.
  • Document standard methodologies for internal sharing.
  • Support R&D, engage in POCs/POVs for future improvements.

Skills

HPC/AI cluster design
Linux/Windows networking
Python scripting
Automation/config management
Networking protocols
Virtualization
Cloud platforms

Education

Bachelor's degree in Computer Science or Engineering

Tools

Slurm
Kubernetes
Jenkins
Ansible
Puppet
Chef
VMware
Hyper-V
KVM
Citrix
Lustre
GPFS
Weka.io
Docker

Job description

NVIDIA is seeking an experienced HPC-AI Engineer to join the Networking Clusters Solutions Infrastructure team. You will design and run large-scale HPC/AI clusters, manage job workloads with Slurm and K8s, and automate deployment via CI/CD pipelines.

You’ll work with researchers and customers to craft workflows and advanced solutions, across Linux/Windows, networking, storage and virtualization layers. Base salary varies by level (approximately 176k–276k USD for Level 4 and 208k–333.5k USD for

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior HPC-AI Cluster Architect (Equity)
Senior HPC-AI Cluster Architect (Equity)

NVIDIA • Santa Clara (CA)

On-site
USD 176,000 - 334,000
Equity
Benefits
Senior HPC-AI Systems Architect (Equity)
Senior HPC-AI Systems Architect (Equity)

Nvidia Corporation in • Santa Clara (CA)

On-site
USD 176,000 - 334,000
Senior HPC AI Cluster Engineer
Senior HPC AI Cluster Engineer

NVIDIA • California (MO)

On-site
USD 176,000 - 334,000
Senior HPC AI Cluster Engineer
Senior HPC AI Cluster Engineer

NVIDIA • Santa Clara (CA)

On-site
USD 176,000 - 334,000
Equity
Benefits
Lead HPC Cluster Engineer for GPU AI Compute
Lead HPC Cluster Engineer for GPU AI Compute

NVIDIA • Santa Clara (CA)

On-site
USD 152,000 - 242,000
Equity
Benefits
Senior HPC Deployment Lead – AI & Data Center Validation
Senior HPC Deployment Lead – AI & Data Center Validation

NVIDIA AI • Holmdel Township (NJ)

On-site
USD 216,000 - 397,000
Equity
Benefits
Senior HPC AI Cluster Engineer
Senior HPC AI Cluster Engineer

Nvidia Corporation in • Santa Clara (CA)

On-site
USD 176,000 - 334,000
Lead AI Infrastructure Solutions Architect for HPC Clusters
Lead AI Infrastructure Solutions Architect for HPC Clusters

NVIDIA • Santa Clara (CA)

On-site
USD 184,000 - 288,000
Senior HPC AI Infra Engineer - Onsite/Remote | Equity
Senior HPC AI Infra Engineer - Onsite/Remote | Equity

NVIDIA • Nashville (TN)

Hybrid
USD 108,000 - 173,000
Equity options
Comprehensive benefits package
Inclusive work environment
Lead AI Infrastructure Solutions Architect for HPC Clusters
Lead AI Infrastructure Solutions Architect for HPC Clusters

NVIDIA • New York (NY)

On-site
USD 184,000 - 357,000