AI/HPC Infrastructure Engineer: GPU Compute & Hybrid Cloud

Saigepartners

San Jose (CA)

Hybrid

USD 120,000 - 180,000

Full time

3 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Saige Partners is seeking an AI/HPC Systems Engineer to build, deploy, and operate GPU-enabled compute infrastructure for AI and HPC workloads across on-premises and cloud. The role emphasizes Linux, GPU/HPC environments, and automation.

You will collaborate with engineering and IT teams to deliver scalable, reliable systems. The successful candidate will work across HPC, GPU computing, cloud infrastructure, and AI platforms, aligning with engineering and R&D needs to support evolving technology

Qualifications

  • Bachelor's degree in Computer Science, Engineering, or a related technical field.
  • 3+ years of hands-on experience in IT infrastructure, cloud engineering, platform engineering, HPC, or a related field.
  • Hands-on experience with Linux-based infrastructure and public cloud platforms such as AWS, Azure, or GCP.
  • Experience deploying, configuring, or operating GPU and/or HPC environments.
  • Experience with workload scheduling or orchestration technologies such as Kubernetes, Slurm, or similar platforms.
  • Experience with infrastructure automation, monitoring, troubleshooting, and performance optimization.
  • Strong understanding of compute, storage, networking, virtualization, and container technologies.
  • Experience supporting AI/ML infrastructure or workloads is a plus.
  • Strong problem-solving, collaboration, and communication skills.
  • Ability to work effectively across engineering, R&D, and IT teams.

Responsibilities

  • GPU/HPC Infrastructure Build, configure, and operate GPU and HPC clusters across compute, storage, and networking environments.
  • Support capacity planning, performance tuning, and infrastructure optimization for AI training, inference, and compute-intensive workloads.
  • Monitor system performance, availability, and resource utilization to ensure reliable operations.
  • Hybrid Cloud Infrastructure Deploy and maintain computing environments across on-premises infrastructure and public cloud platforms.
  • Support infrastructure modernization, expansion, and scaling initiatives for HPC and AI workloads.
  • Help evaluate and implement solutions that improve scalability, reliability, and cost efficiency.
  • Automation & Observability Implement infrastructure-as-code and automated provisioning solutions.
  • Develop and maintain monitoring, logging, alerting, and observability capabilities.
  • Automate routine infrastructure tasks and identify opportunities to improve resource utilization and operational efficiency.
  • AI Platform Support Deploy, integrate, and support LLM APIs, coding assistants, and AI/agent platforms used by engineering teams.
  • Assist with the infrastructure requirements and operational support of AI/ML workloads.
  • Collaborate with engineering teams to ensure AI platforms are reliable, accessible, and scalable.
  • Infrastructure Operations & Troubleshooting Troubleshoot and resolve infrastructure, networking, compute, storage, and platform issues.
  • Support day-to-day IT and infrastructure operations across engineering environments.
  • Develop and maintain technical documentation, standards, procedures, and operational runbooks.
  • Collaborate with engineering, IT, and other stakeholders to deliver reliable infrastructure solutions.

Skills

IT infrastructure
Cloud engineering
Platform engineering
HPC
Linux infrastructure
GPU/HPC environments
Automation & monitoring
AI/ML infrastructure
Problem solving
Cross-team collaboration

Education

Bachelor's degree in Computer Science, Engineering, or a related field

Tools

Kubernetes
Slurm
AWS
Azure
GCP

Job description

Saige Partners is seeking an AI/HPC Systems Engineer to build, deploy, and operate GPU-enabled compute infrastructure for AI and HPC workloads across on-premises and cloud. The role emphasizes Linux, GPU/HPC environments, and automation.

You will collaborate with engineering and IT teams to deliver scalable, reliable systems. The successful candidate will work across HPC, GPU computing, cloud infrastructure, and AI platforms, aligning with engineering and R&D needs to support evolving technology

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI/HPC Systems Engineer
AI/HPC Systems Engineer

Saigepartners • San Jose (CA)

Hybrid
USD 120,000 - 180,000
AI/HPC Systems Engineer — Hybrid Cloud & GPUs
AI/HPC Systems Engineer — Hybrid Cloud & GPUs

Protingent • San Jose (CA)

On-site
USD 110,000 - 124,000
Insurance plan options (HDHP/POS)
Pre-tax commuter benefits
401k plan
+1
AI/HPC Systems Engineer – Hybrid Cloud GPU Infra
AI/HPC Systems Engineer – Hybrid Cloud GPU Infra

Norland Group • San Jose (CA)

On-site
USD 103,000 - 117,000
Hybrid AI HPC Infrastructure Engineer (GPU/ML)
Hybrid AI HPC Infrastructure Engineer (GPU/ML)

Analysis Group, Inc. • Boston (MA)

On-site
USD 150,000 - 170,000
Discretionary annual bonus
Benefits package
Hybrid HPC Systems Architect - GPU Cloud for AI
Hybrid HPC Systems Architect - GPU Cloud for AI

The Consensus • San Jose (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Cash compensation
Equity compensation
Health, dental and vision coverage
+1
Hybrid GPU Data Center Engineer: Automation & AI Infra
Hybrid GPU Data Center Engineer: Automation & AI Infra

Blue Signal Search • United States

Hybrid
USD 120,000 - 180,000
Competitive compensation
Equity opportunity
Comprehensive benefits
AI Systems Engineer: HPC & GPU Clusters
AI Systems Engineer: HPC & GPU Clusters

Advanced Micro Devices, Inc. • San Jose (CA)

On-site
USD 180,000 - 260,000
AI Systems Engineer: HPC & GPU Clusters
AI Systems Engineer: HPC & GPU Clusters

AMD • San Jose (CA)

On-site
USD 180,000 - 260,000
AMD benefits
Senior HPC Cloud Engineer — GPU/InfiniBand AI Infra
Senior HPC Cloud Engineer — GPU/InfiniBand AI Infra

Jobgether • Germany (OH)

On-site
USD 81,000 - 105,000
Career development opportunities
Flexible working arrangements
Collaborative engineering environment
+1
AI Compute Sales Lead - Hybrid GPU & Cloud Solutions
AI Compute Sales Lead - Hybrid GPU & Cloud Solutions

Parallel Works • Chicago (IL)

On-site
USD 120,000 - 180,000
Medical, vision, and dental coverage
401(k) with company match
Generous paid vacation and sick time
+1