Senior HPC-AI Systems Architect (Equity)

Nvidia Corporation in

Santa Clara (CA)

On-site

USD 176,000 - 334,000

Full time

8 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

NVIDIA is seeking an experienced Senior HPC-AI Cluster Engineer to join the Networking Clusters Solutions Infrastructure team in Santa Clara. You will design and scale large HPC/AI clusters, manage job scheduling, and develop CI/CD pipelines for fast, reliable AI workloads.

You will work across Linux/Windows environments, implement automation, monitor infrastructure, and support R&D initiatives to push the boundaries of GPU computing and AI systems.

Qualifications

  • A degree in Computer Science, Engineering, or a related field (or equivalent experience) and 8+ years of experience.
  • Knowledge of HPC and AI solution technologies from CPU's and GPU's to high speed interconnects and supporting software.
  • Experience with job scheduling workloads and orchestration tools such as Slurm, K8s.
  • Excellent knowledge of Windows and Linux networking and internals and OS level security.

Responsibilities

  • Design, implement and maintain large scale HPC/AI clusters with monitoring, logging and alerting.
  • Manage Linux job/workload schedules and orchestration tools
  • Develop and maintain continuous integration and delivery pipelines
  • Develop tooling to automate deployment and management of large-scale infrastructure environments
  • Deploy monitoring solutions for the servers, network and storage
  • Perform troubleshooting bottom up from bare metal to application level
  • Develop, redefine and document standard methodologies for internal use
  • Support R&D activities and engage in POCs/POVs for future improvements

Skills

HPC/AI cluster design
Linux system administration
Job scheduling (Slurm, K8s)
Networking (InfiniBand, Ethernet)
Python scripting
Automation tools (Jenkins, Ansible, Pu

Education

Bachelor's degree in Computer Science or Engineering

Tools

Slurm
Kubernetes
VMware/Hyper-V/KVM
GPFS/Lustre

Job description

NVIDIA is seeking an experienced Senior HPC-AI Cluster Engineer to join the Networking Clusters Solutions Infrastructure team in Santa Clara. You will design and scale large HPC/AI clusters, manage job scheduling, and develop CI/CD pipelines for fast, reliable AI workloads.

You will work across Linux/Windows environments, implement automation, monitor infrastructure, and support R&D initiatives to push the boundaries of GPU computing and AI systems.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior HPC-AI Cluster Architect (Equity)
Senior HPC-AI Cluster Architect (Equity)

NVIDIA • Santa Clara (CA)

On-site
USD 176,000 - 334,000
Equity
Benefits
Senior HPC-AI Cluster Architect (Equity)
Senior HPC-AI Cluster Architect (Equity)

NVIDIA • California (MO)

On-site
USD 176,000 - 334,000
Senior HPC AI Cluster Engineer
Senior HPC AI Cluster Engineer

NVIDIA • California (MO)

On-site
USD 176,000 - 334,000
Senior HPC Architect - GPU Compute, Equity Eligible
Senior HPC Architect - GPU Compute, Equity Eligible

NVIDIA • California (MO)

On-site
USD 184,000 - 288,000
Equity
Inclusive work environment
Comprehensive benefits
Senior AI/HPC Solutions Architect – Linux & Networking
Senior AI/HPC Solutions Architect – Linux & Networking

Nvidia Corporation • Santa Clara (CA)

On-site
USD 148,000 - 236,000
Senior AI Factory Architect — Multi-GPU HPC, NCCL, Equity
Senior AI Factory Architect — Multi-GPU HPC, NCCL, Equity

NVIDIA • California (MO)

On-site
USD 152,000 - 288,000
Equity
Benefits
Senior AI Infrastructure Architect for Enterprise ISVs
Senior AI Infrastructure Architect for Enterprise ISVs

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 184,000 - 288,000
Equity
Comprehensive benefits
Performance bonus
Senior AI Compute Engineer - HPC Infra & Linux
Senior AI Compute Engineer - HPC Infra & Linux

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 148,000 - 288,000
Senior AI Compute Architect for Enterprise Data Centers
Senior AI Compute Architect for Enterprise Data Centers

AIToolboard • United States

On-site
USD 184,000 - 357,000
Equity
Benefits
Senior HPC Architect: At-Scale GPU Deployment & Automation
Senior HPC Architect: At-Scale GPU Deployment & Automation

NVIDIA • New Mexico

On-site
USD 184,000 - 288,000