AI/ML HPC Cluster Engineer — Scale GPU-Accelerated Infra

NVIDIA

Colorado

On-site

USD 124,000 - 195,500

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

A leading technology company located in Colorado is seeking to fill a critical role in managing AI superclusters. The successful candidate will support day-to-day operations of both on-premises and multi-cloud AI/HPC clusters, ensuring optimal performance and user satisfaction. The role requires a Bachelor's degree and at least 2 years of experience with compute infrastructure and AI job schedulers. The base salary ranges from 124,000 to 195,500 USD, accompanied by equity and benefits. Applications accepted until February 24, 2026.

Qualifications

  • Minimum 2+ years of experience administering multi-node compute infrastructure.
  • Background in managing AI/HPC job schedulers.
  • Proficient in Centos/RHEL and/or Ubuntu Linux distributions.

Responsibilities

  • Support day-to-day operations of AI/HPC clusters.
  • Directly administer internal research clusters.
  • Develop and improve the ecosystem around GPU-accelerated computing.

Skills

Administration of multi-node compute infrastructure
Working with AI/HPC job schedulers
Proficient in Centos/RHEL, Ubuntu
Cluster configuration management tools
Container technologies
Python programming
Bash scripting

Education

Bachelor’s degree in Computer Science, Electrical Engineering or related field

Tools

Ansible
Docker
Kubernetes

Job description

A leading technology company located in Colorado is seeking to fill a critical role in managing AI superclusters. The successful candidate will support day-to-day operations of both on-premises and multi-cloud AI/HPC clusters, ensuring optimal performance and user satisfaction. The role requires a Bachelor's degree and at least 2 years of experience with compute infrastructure and AI job schedulers. The base salary ranges from 124,000 to 195,500 USD, accompanied by equity and benefits. Applications accepted until February 24, 2026.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI GPU Infra Engineer — Performance & Scale
Senior AI GPU Infra Engineer — Performance & Scale

Summit Group Solutions, LLC • United States

On-site
USD 150,000 - 350,000
AI and ML HPC Cluster Engineer, AI and ML HPC Cluster Engineer
AI and ML HPC Cluster Engineer, AI and ML HPC Cluster Engineer

NVIDIA • Colorado

On-site
USD 124,000 - 196,000
Software Engineer, Scalable AI HPC Platform
Software Engineer, Scalable AI HPC Platform

Magic • United States

On-site
USD 200,000 - 550,000
Lead HPC Cluster Engineer for GPU AI Compute
Lead HPC Cluster Engineer for GPU AI Compute

NVIDIA • Santa Clara (CA)

On-site
USD 152,000 - 242,000
Equity
Benefits
Senior AI Infra Engineer: GPU Clusters & Kubernetes
Senior AI Infra Engineer: GPU Clusters & Kubernetes

Intelliswift - An LTTS Company • Sunnyvale (CA)

On-site
USD 120,000 - 150,000
Lead AI Infrastructure Solutions Architect for HPC Clusters
Lead AI Infrastructure Solutions Architect for HPC Clusters

NVIDIA • Santa Clara (CA)

On-site
USD 184,000 - 288,000
Senior AI Infra Engineer — HPC & Scheduler
Senior AI Infra Engineer — HPC & Scheduler

Ai2 • Seattle (WA)

On-site
USD 126,000 - 189,000
Medical, dental, and vision insurance
401(k) plan enrollment
Monthly stipends for commuting and fitness
+1
Staff Compute Infra Engineer - GPU & AI Systems
Staff Compute Infra Engineer - GPU & AI Systems

xAI • Palo Alto (CA)

On-site
USD 180,000 - 440,000
Senior AI Performance & Efficiency Engineer - GPU Clusters
Senior AI Performance & Efficiency Engineer - GPU Clusters

NVIDIA • Santa Clara (CA)

On-site
USD 152,000 - 241,500
Equity
Comprehensive benefits package
Senior HPC-AI Cluster Architect (Equity)
Senior HPC-AI Cluster Architect (Equity)

NVIDIA • California (MO)

On-site
USD 176,000 - 334,000