AI/HPC System Engineer

Norland Group

San Jose (CA)

On-site

USD 103,000 - 117,000

Full time

4 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Norland Group is seeking an AI/HPC System Engineer to build and operate GPU/HPC clusters across on-premise and cloud environments. You will deploy, automate, and maintain compute infrastructure to support our AI and R&D workloads.

The role requires 3+ years in IT infrastructure or HPC, Linux and public cloud experience, and familiarity with Kubernetes or Slurm for workload orchestration. Collaboration with engineering teams is essential.

Qualifications

  • Bachelor's degree in Computer Science, Engineering, or related field.
  • 3+ years of hands-on IT infrastructure, cloud, platform engineering, or HPC experience.
  • Hands-on Linux-based infrastructure and public cloud environments (AWS/Azure/GCP).
  • Experience deploying or operating GPU/HPC environments with workload scheduling/orchestration (Kubernetes or Slurm).
  • Experience with infrastructure automation, monitoring, troubleshooting and performance optimization.

Responsibilities

  • Build, configure, and operate GPU and HPC clusters across compute, storage, and networking.
  • Deploy and maintain hybrid cloud compute environments (on-premise and public cloud).
  • Implement infrastructure-as-code, provisioning automation, monitoring, and alerting.
  • Deploy, integrate, and support LLM APIs, coding assistants, and AI agent platforms used by internal teams.
  • Troubleshoot infrastructure issues, document standards/runbooks, and support day-to-day IT operations.

Skills

Linux fundamentals
Cloud platforms
GPU/HPC experience
Kubernetes
Slurm
Automation tooling
Monitoring observability
Collaboration

Education

Bachelor's degree in CS/Engineering

Tools

Kubernetes
Slurm
AWS
Azure
GCP

Job description

This position requires - Clear Background, Drug Test, and Education Check. Must be authorized to work in the US for any employer without Sponsorship. (Principal Only! No Corp to Corp)

Contract Duration: 6 months contract

Description

We are hiring an AI/HPC System Engineer to build and operate the compute infrastructure supporting our HPC and AI development workloads. This role deploys, automates, and maintains GPU clusters across on-premise and cloud environments, delivering reliable, scalable, and cost-efficient compute for engineering and R&D teams.

Responsibilities
  • GPU/HPC infrastructure: Build, configure, and operate GPU and HPC clusters across compute, storage, and networking; support capacity planning, performance tuning, and optimization for AI training, inference, and compute-intensive workloads
  • Hybrid cloud infrastructure: Deploy and maintain compute environments spanning on-premise and public cloud, and contribute to modernization and scaling initiatives for HPC/AI infrastructure
  • Automation and observability: Implement infrastructure-as-code, provisioning automation, monitoring, and alerting, and drive improvements in resource utilization and efficiency
  • AI platform support: Deploy, integrate, and support LLM APIs, coding assistants, and AI/agent platforms used by internal engineering teams
  • Operations and collaboration: Troubleshoot and resolve infrastructure issues, document standards and runbooks, and work with relevant stakeholders to support day-to-day IT operations
Requirements
  • Bachelor's degree in Computer Science, Engineering, or a related technical field
  • 3+ years of hands-on experience in IT infrastructure, cloud, platform engineering, or HPC
  • Hands-on experience with Linux-based infrastructure and public cloud environments such as AWS, Azure, or GCP
  • Experience deploying or operating GPU/HPC environments, including workload scheduling or orchestration platforms such as Kubernetes or Slurm
  • Experience with infrastructure automation, monitoring, troubleshooting, and performance optimization
  • Solid understanding of compute, storage, networking, and container technologies; experience with AI/ML infrastructure or workloads is a plus
  • Strong collaboration and communication skills, with the ability to work across engineering and IT teams

We encourage Minorities, Women, Protected Veterans and Disabled individuals to apply for all positions that they may be qualified for. We maintain a drug-free workplace and perform pre-employment substance abuse testing and background checks.

Position Title: AI/HPC System Engineer

Location: San Jose, CA

Pay Rate: $75-$85

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI/HPC System Engineer
AI/HPC System Engineer

Protingent • San Jose (CA)

On-site
USD 110,000 - 124,000
Insurance plan options (HDHP/POS)
Pre-tax commuter benefits
401k plan
+1
AI Systems Engineer - HPC
AI Systems Engineer - HPC

Advanced Micro Devices, Inc. • San Jose (CA)

On-site
USD 180,000 - 260,000
AI Systems Engineer - HPC
AI Systems Engineer - HPC

Advanced Micro Devices • San Jose (CA), Northern (KY)

Hybrid
USD 140,000 - 190,000
AI/HPC Cluster Design Engineer
AI/HPC Cluster Design Engineer

Advanced Micro Devices, Inc. • Austin (TX)

On-site
USD 120,000 - 190,000
AMD benefits at a glance
AI/HPC Cluster Design Engineer
AI/HPC Cluster Design Engineer

Advanced Micro Devices • Austin (TX)

On-site
USD 120,000 - 180,000
AMD benefits
Senior HPC AI Cluster Engineer
Senior HPC AI Cluster Engineer

NVIDIA • California (MO)

On-site
USD 176,000 - 334,000
Software Engineer, Compute Infrastructure
Software Engineer, Compute Infrastructure

OpenAI • Los Angeles (CA)

On-site
USD 230,000 - 405,000
Equity
Flexible work environment
Health benefits
AI/HPC Systems Engineer — Hybrid Cloud & GPUs
AI/HPC Systems Engineer — Hybrid Cloud & GPUs

Protingent • San Jose (CA)

On-site
USD 110,000 - 124,000
Insurance plan options (HDHP/POS)
Pre-tax commuter benefits
401k plan
+1
Senior HPC AI Cluster Engineer
Senior HPC AI Cluster Engineer

Nvidia Corporation in • Santa Clara (CA)

On-site
USD 176,000 - 334,000
AI/HPC Cluster Design Engineer
AI/HPC Cluster Design Engineer

AMD • Austin (TX)

On-site
USD 140,000 - 200,000