AI/HPC Systems Engineer

Saigepartners

San Jose (CA)

Hybrid

USD 120,000 - 180,000

Full time

3 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Saige Partners is seeking an AI/HPC Systems Engineer to build, deploy, and operate GPU-enabled compute infrastructure for AI and HPC workloads across on-premises and cloud. The role emphasizes Linux, GPU/HPC environments, and automation.

You will collaborate with engineering and IT teams to deliver scalable, reliable systems. The successful candidate will work across HPC, GPU computing, cloud infrastructure, and AI platforms, aligning with engineering and R&D needs to support evolving technology

Qualifications

  • Bachelor's degree in Computer Science, Engineering, or a related technical field.
  • 3+ years of hands-on experience in IT infrastructure, cloud engineering, platform engineering, HPC, or a related field.
  • Hands-on experience with Linux-based infrastructure and public cloud platforms such as AWS, Azure, or GCP.
  • Experience deploying, configuring, or operating GPU and/or HPC environments.
  • Experience with workload scheduling or orchestration technologies such as Kubernetes, Slurm, or similar platforms.
  • Experience with infrastructure automation, monitoring, troubleshooting, and performance optimization.
  • Strong understanding of compute, storage, networking, virtualization, and container technologies.
  • Experience supporting AI/ML infrastructure or workloads is a plus.
  • Strong problem-solving, collaboration, and communication skills.
  • Ability to work effectively across engineering, R&D, and IT teams.

Responsibilities

  • GPU/HPC Infrastructure Build, configure, and operate GPU and HPC clusters across compute, storage, and networking environments.
  • Support capacity planning, performance tuning, and infrastructure optimization for AI training, inference, and compute-intensive workloads.
  • Monitor system performance, availability, and resource utilization to ensure reliable operations.
  • Hybrid Cloud Infrastructure Deploy and maintain computing environments across on-premises infrastructure and public cloud platforms.
  • Support infrastructure modernization, expansion, and scaling initiatives for HPC and AI workloads.
  • Help evaluate and implement solutions that improve scalability, reliability, and cost efficiency.
  • Automation & Observability Implement infrastructure-as-code and automated provisioning solutions.
  • Develop and maintain monitoring, logging, alerting, and observability capabilities.
  • Automate routine infrastructure tasks and identify opportunities to improve resource utilization and operational efficiency.
  • AI Platform Support Deploy, integrate, and support LLM APIs, coding assistants, and AI/agent platforms used by engineering teams.
  • Assist with the infrastructure requirements and operational support of AI/ML workloads.
  • Collaborate with engineering teams to ensure AI platforms are reliable, accessible, and scalable.
  • Infrastructure Operations & Troubleshooting Troubleshoot and resolve infrastructure, networking, compute, storage, and platform issues.
  • Support day-to-day IT and infrastructure operations across engineering environments.
  • Develop and maintain technical documentation, standards, procedures, and operational runbooks.
  • Collaborate with engineering, IT, and other stakeholders to deliver reliable infrastructure solutions.

Skills

IT infrastructure
Cloud engineering
Platform engineering
HPC
Linux infrastructure
GPU/HPC environments
Automation & monitoring
AI/ML infrastructure
Problem solving
Cross-team collaboration

Education

Bachelor's degree in Computer Science, Engineering, or a related field

Tools

Kubernetes
Slurm
AWS
Azure
GCP

Job description

We strive to be Your Future, Your Solution to accelerate your career! This is a W2 contract position and is not eligible for C2C or W2 referral candidates.

AI/HPC Systems Engineer Position Overview

We are seeking an AI/HPC Systems Engineer to build, deploy, and operate the compute infrastructure supporting high-performance computing and AI development workloads. This role will focus on deploying, automating, and maintaining GPU-enabled environments across on-premises and cloud platforms while delivering reliable, scalable, and cost-effective computing resources for engineering and R&D teams. The ideal candidate brings hands-on experience with Linux infrastructure, GPU/HPC environments, cloud platforms, automation, and modern AI/ML infrastructure.

Key Responsibilities
  • GPU/HPC Infrastructure Build, configure, and operate GPU and HPC clusters across compute, storage, and networking environments.
  • Support capacity planning, performance tuning, and infrastructure optimization for AI training, inference, and compute-intensive workloads.
  • Monitor system performance, availability, and resource utilization to ensure reliable operations.
  • Hybrid Cloud Infrastructure Deploy and maintain computing environments across on-premises infrastructure and public cloud platforms.
  • Support infrastructure modernization, expansion, and scaling initiatives for HPC and AI workloads.
  • Help evaluate and implement solutions that improve scalability, reliability, and cost efficiency.
  • Automation & Observability Implement infrastructure-as-code and automated provisioning solutions.
  • Develop and maintain monitoring, logging, alerting, and observability capabilities.
  • Automate routine infrastructure tasks and identify opportunities to improve resource utilization and operational efficiency.
  • AI Platform Support Deploy, integrate, and support LLM APIs, coding assistants, and AI/agent platforms used by engineering teams.
  • Assist with the infrastructure requirements and operational support of AI/ML workloads.
  • Collaborate with engineering teams to ensure AI platforms are reliable, accessible, and scalable.
  • Infrastructure Operations & Troubleshooting Troubleshoot and resolve infrastructure, networking, compute, storage, and platform issues.
  • Support day-to-day IT and infrastructure operations across engineering environments.
  • Develop and maintain technical documentation, standards, procedures, and operational runbooks.
  • Collaborate with engineering, IT, and other stakeholders to deliver reliable infrastructure solutions.
Qualifications
  • Bachelor's degree in Computer Science, Engineering, or a related technical field.
  • 3+ years of hands-on experience in IT infrastructure, cloud engineering, platform engineering, HPC, or a related field.
  • Hands-on experience with Linux-based infrastructure and public cloud platforms such as AWS, Azure, or GCP.
  • Experience deploying, configuring, or operating GPU and/or HPC environments.
  • Experience with workload scheduling or orchestration technologies such as Kubernetes, Slurm, or similar platforms.
  • Experience with infrastructure automation, monitoring, troubleshooting, and performance optimization.
  • Strong understanding of compute, storage, networking, virtualization, and container technologies.
  • Experience supporting AI/ML infrastructure or workloads is a plus.
  • Strong problem-solving, collaboration, and communication skills.
  • Ability to work effectively across engineering, R&D, and IT teams.

What You'll Bring The successful candidate will be a hands-on infrastructure professional who enjoys solving complex technical challenges and building reliable systems at scale. You should be comfortable working across HPC, GPU computing, cloud infrastructure, Linux, automation, containers, and AI platforms, while partnering closely with engineering and IT teams to support evolving technology needs.

Learn more about Saige Partners on Facebook or LinkedIn. Saige Partners, one of the fastest growing technology and talent companies in the Midwest, believes in people with a passion to help them succeed. We are in the business of helping professionals Build Careers, Not Jobs. Saige Partners believes employees are the most valuable asset to building a thriving and successful company culture, which is why we offer a benefit package and convenient weekly payment solutions that helps our employees stay healthy and maintain a positive work/life balance.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI/HPC Infrastructure Engineer: GPU Compute & Hybrid Cloud
AI/HPC Infrastructure Engineer: GPU Compute & Hybrid Cloud

Saigepartners • San Jose (CA)

Hybrid
USD 120,000 - 180,000
HPC AI Systems Administrator
HPC AI Systems Administrator

MRE Consulting • Houston (TX)

On-site
USD 95,000 - 140,000
AI/HPC System Engineer
AI/HPC System Engineer

Norland Group • San Jose (CA)

On-site
USD 103,000 - 117,000
AI Infra/HPC Engineer
AI Infra/HPC Engineer

Blue Signal Search • San Francisco (CA)

On-site
USD 180,000 - 240,000
Annual bonus
Equity participation
Comprehensive benefits
+1
Senior Solutions Engineer, AI Infrastructure
Senior Solutions Engineer, AI Infrastructure

VAST Data • New York (NY)

On-site
USD 150,000 - 200,000
HPC Solution Architect - AI Infrastructure
HPC Solution Architect - AI Infrastructure

Hamilton Barnes Associates Limited • San Francisco (CA)

On-site
USD 270,000 - 330,000
Equity
Healthcare + Dental + Vision
401(k)
+1
AI Platform Engineer
AI Platform Engineer

Park Place Technologies in • Highland Heights (OH)

On-site
USD 90,000 - 130,000
HPC/AI Technical Solution Engineer
HPC/AI Technical Solution Engineer

VC5 Consulting • Houston (TX)

On-site
USD 120,000 - 180,000
Head of AI Data Center Infrastructure Platforms and Software
Head of AI Data Center Infrastructure Platforms and Software

Summit Group Solutions, LLC • United States

On-site
USD 150,000 - 350,000
Data Center Compute Engineer
Data Center Compute Engineer

Blue Signal Search • United States

Hybrid
USD 120,000 - 180,000
Competitive compensation
Equity opportunity
Comprehensive benefits