Senior HPC-AI Cluster Architect (Equity)

NVIDIA

Santa Clara (CA)

On-site

USD 176,000 - 334,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Equity
Benefits

Job summary

NVIDIA is seeking an experienced HPC-AI Engineer to join the Networking Clusters Solutions Infrastructure team. You will design, implement, and maintain large-scale HPC/AI clusters with monitoring and automation, work with Linux/Windows networking, and develop deployment pipelines for cutting-edge AI workloads.

Responsibilities include tuning, scheduling, and automating compute workflows, plus collaboration with researchers and customers to deliver scalable, high-performance solutions.

Qualifications

  • Degree in Computer Science, Engineering, or related field or equivalent experience.
  • 8+ years of experience in HPC/AI infrastructure.
  • Experience with job scheduling/workloads and orchestration tools such as Slurm, Kubernetes.
  • Excellent knowledge of Windows and Linux networking and internals, firewalls, and common protocols.
  • Experience with storage solutions (Lustre, GPFS, Weka.io) and emerging technologies.
  • Python programming and Bash scripting.
  • Familiarity with automation/configuration tools (Jenkins, Ansible, Puppet/Chef).
  • Deep knowledge of Networking Protocols (InfiniBand, Ethernet).
  • Experience with virtualization (VMware/Hyper-V/KVM/Citrix).
  • Familiarity with cloud platforms (AWS, Azure, Google Cloud).

Responsibilities

  • Design, implement and maintain large-scale HPC/AI clusters with monitoring, logging, and alerting.
  • Manage Linux job/workload schedules and orchestration tools.
  • Develop and maintain CI/CD pipelines.
  • Develop tooling to automate deployment and management of large-scale infra.
  • Deploy monitoring solutions for servers, network, and storage.
  • Troubleshoot from bare metal to application level.
  • Document standard methodologies for internal teams.
  • Support R&D activities and engage in POCs/POVs for improvements.

Skills

HPC/AI systems design
Linux
Job scheduling (Slurm, K8s)
Networking (InfiniBand, Ethernet)
Python
Shell scripting
CI/CD (Jenkins, Ansible)
Virtualization (VMware/Hyper-V/KVM)
Cloud platforms (AWS/Azure/GCP)

Education

Bachelor's degree in CS/Engineering or related field

Tools

Jenkins
Ansible
Kubernetes
Puppet/Chef

Job description

NVIDIA is seeking an experienced HPC-AI Engineer to join the Networking Clusters Solutions Infrastructure team. You will design, implement, and maintain large-scale HPC/AI clusters with monitoring and automation, work with Linux/Windows networking, and develop deployment pipelines for cutting-edge AI workloads.

Responsibilities include tuning, scheduling, and automating compute workflows, plus collaboration with researchers and customers to deliver scalable, high-performance solutions.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior HPC-AI Cluster Architect (Equity)
Senior HPC-AI Cluster Architect (Equity)

NVIDIA • California (MO)

On-site
USD 176,000 - 334,000
Senior AI/HPC Networking Solutions Architect
Senior AI/HPC Networking Solutions Architect

NVIDIA AI • Indiana (PA)

On-site
USD 140,000 - 210,000
Senior HPC AI Cluster Engineer
Senior HPC AI Cluster Engineer

NVIDIA • California (MO)

On-site
USD 176,000 - 334,000
Senior HPC Architect: Large-Scale GPU AI Infra (Equity)
Senior HPC Architect: Large-Scale GPU AI Infra (Equity)

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Equity
Comprehensive benefits
Senior AI Compute Architect for Enterprise Data Centers
Senior AI Compute Architect for Enterprise Data Centers

AIToolboard • United States

On-site
USD 184,000 - 357,000
Equity
Benefits
Senior AI/HPC Solutions Architect – Linux & Networking
Senior AI/HPC Solutions Architect – Linux & Networking

Nvidia Corporation • Santa Clara (CA)

On-site
USD 148,000 - 236,000
Senior Solutions Architect - AI Cluster Networking Design
Senior Solutions Architect - AI Cluster Networking Design

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Equity
Benefits
Senior AI Factory Architect — Multi-GPU HPC, NCCL, Equity
Senior AI Factory Architect — Multi-GPU HPC, NCCL, Equity

NVIDIA • California (MO)

On-site
USD 152,000 - 288,000
Equity
Benefits
Lead AI Infrastructure Solutions Architect for HPC Clusters
Lead AI Infrastructure Solutions Architect for HPC Clusters

NVIDIA • New York (NY)

On-site
USD 184,000 - 357,000
Senior HPC Deployment & Validation Lead Equity
Senior HPC Deployment & Validation Lead Equity

NVIDIA • South Carolina

On-site
USD 230,000 - 360,000