AI/HPC Systems Engineer — GPU Cluster & Hybrid Cloud

Saige Partners

San Jose (CA)

On-site

USD 140,000 - 190,000

Full time

13 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Saige Partners seeks an AI/HPC Systems Engineer to build, deploy, and operate GPU-enabled compute infrastructure across on-premises and cloud platforms. You will automate provisioning, monitor performance, and support AI/ML workloads for engineering and R&D teams.

The role focuses on GPU/HPC clusters, hybrid cloud, and observability, collaborating with IT and engineering to deliver reliable systems at scale.

Qualifications

  • Bachelor's degree in Computer Science, Engineering, or a related technical field.
  • 3 years of hands-on IT infrastructure, cloud, HPC, or related experience.
  • Experience with Linux-based infrastructure and public clouds (AWS, Azure, or GCP).
  • Experience deploying, configuring, or operating GPU and/or HPC environments.
  • Experience with workload scheduling or orchestration such as Kubernetes, Slurm, or similar.
  • Experience with infrastructure automation, monitoring, troubleshooting, and performance optimization.
  • Strong understanding of compute, storage, networking, virtualization, and container technologies.
  • Experience supporting AI/ML infrastructure or workloads is a plus.
  • Strong problem-solving, collaboration, and communication skills.
  • Ability to work effectively across engineering, R&D, and IT teams.

Responsibilities

  • Build, configure, and operate GPU and HPC clusters across compute, storage, and networking environments.
  • Support capacity planning, performance tuning, and infrastructure optimization for AI training, inference, and compute-intensive workloads.
  • Monitor system performance, availability, and resource utilization to ensure reliable operations.
  • Deploy and maintain computing environments across on-premises infrastructure and public cloud platforms.
  • Support infrastructure modernization, expansion, and scaling initiatives for HPC and AI workloads.
  • Help evaluate and implement solutions that improve scalability, reliability, and cost efficiency.
  • Implement infrastructure-as-code and automated provisioning solutions.
  • Develop and maintain monitoring, logging, alerting, and observability capabilities.
  • Automate routine infrastructure tasks and identify opportunities to improve resource utilization and operational efficiency.
  • Deploy, integrate, and support LLM APIs, coding assistants, and AI/agent platforms used by engineering teams.
  • Assist with the infrastructure requirements and operational support of AI/ML workloads.
  • Collaborate with engineering teams to ensure AI platforms are reliable, accessible, and scalable.
  • Troubleshoot and resolve infrastructure, networking, compute, storage, and platform issues.
  • Develop and maintain technical documentation, standards, procedures, and operational runbooks.
  • Collaborate with engineering, IT, and other stakeholders to deliver reliable infrastructure solutions.

Skills

GPU HPC infra
Linux systems
Public cloud (AWS)
Kubernetes
Slurm
Automation & monitoring
AI/ML infra
Collaborative teamwork
Communication skills

Education

Bachelor's degree in Computer Science, Engineering, or a related technical field

Tools

Terraform
Ansible
Docker
Monitoring tools

Job description

Saige Partners seeks an AI/HPC Systems Engineer to build, deploy, and operate GPU-enabled compute infrastructure across on-premises and cloud platforms. You will automate provisioning, monitor performance, and support AI/ML workloads for engineering and R&D teams.

The role focuses on GPU/HPC clusters, hybrid cloud, and observability, collaborating with IT and engineering to deliver reliable systems at scale.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI/HPC Infrastructure Engineer: GPU Compute & Hybrid Cloud
AI/HPC Infrastructure Engineer: GPU Compute & Hybrid Cloud

Saigepartners • San Jose (CA)

Hybrid
USD 120,000 - 180,000
AI/HPC Systems Engineer
AI/HPC Systems Engineer

Saigepartners • San Jose (CA)

Hybrid
USD 120,000 - 180,000
AI/HPC Systems Engineer
AI/HPC Systems Engineer

Saige Partners • San Jose (CA)

On-site
USD 140,000 - 190,000
AI/HPC Systems Engineer – Hybrid Cloud GPU Infra
AI/HPC Systems Engineer – Hybrid Cloud GPU Infra

Norland Group • San Jose (CA)

On-site
USD 103,000 - 117,000
Hybrid AI HPC Infrastructure Engineer (GPU/ML)
Hybrid AI HPC Infrastructure Engineer (GPU/ML)

Analysis Group, Inc. • Boston (MA)

On-site
USD 150,000 - 170,000
Discretionary annual bonus
Benefits package
AI Infrastructure Ops Engineer: GPU Clusters, Hybrid Cloud
AI Infrastructure Ops Engineer: GPU Clusters, Hybrid Cloud

Accenture • Miami (FL)

On-site
USD 87,000 - 266,000
AI Systems Engineer: HPC & GPU Clusters
AI Systems Engineer: HPC & GPU Clusters

Advanced Micro Devices, Inc. • San Jose (CA)

On-site
USD 180,000 - 260,000
Hybrid HPC Systems Architect - GPU Cloud for AI
Hybrid HPC Systems Architect - GPU Cloud for AI

The Consensus • San Jose (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Cash compensation
Equity compensation
Health, dental and vision coverage
+1
AI Systems Engineer: HPC & GPU Clusters
AI Systems Engineer: HPC & GPU Clusters

AMD • San Jose (CA)

On-site
USD 180,000 - 260,000
AMD benefits
AI Infrastructure Engineer: GPU Clusters & Hybrid Cloud
AI Infrastructure Engineer: GPU Clusters & Hybrid Cloud

Accenture • Redmond (WA)

On-site
USD 101,000 - 245,000