AI HPC Systems Engineer: GPU Clusters & ML Platforms

AMD

San Jose (CA)

On-site

USD 140,000 - 210,000

Full time

7 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

AMD IT compute platforms team seeks an AI Systems Engineer to design, deploy, and operate HPC infrastructure, GPU clusters, and AI workload schedulers. You will drive scalable, high-performance AI services on AMD hardware.

You will implement automation, monitor systems, and collaborate with global teams, applying SLURM, Kubernetes, and industry best practices to optimize performance and reliability.

Qualifications

  • Experience with large-scale distributed AI/HPC systems.
  • Deploying GPU clusters and AI workloads with reliability.
  • Proficiency in Linux, Python, and shell scripting.

Responsibilities

  • Develop, implement, and maintain GPU-based clusters with optimal performance.
  • Administer ML/AI platforms, distributed ML services, LLMs, and AI inferencing.
  • Automate system provisioning and end-to-end cluster management.
  • Collaborate with cross-functional teams to address AI infrastructure needs.
  • Monitor performance and ensure security and best practices.

Skills

Python
AI/ML
HPC
Kubernetes
SLURM
GPU clusters
Automation
Terraform
Prometheus
Grafana
Ansible
Shell scripting

Education

Bachelor's or master's in CS/CE

Tools

Kubernetes
SLURM
KVM
Ubuntu
Python
Shell
GPU drivers
RoCEv2
400G networking

Job description

AMD IT compute platforms team seeks an AI Systems Engineer to design, deploy, and operate HPC infrastructure, GPU clusters, and AI workload schedulers. You will drive scalable, high-performance AI services on AMD hardware.

You will implement automation, monitor systems, and collaborate with global teams, applying SLURM, Kubernetes, and industry best practices to optimize performance and reliability.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Systems Engineer: HPC & GPU Clusters
AI Systems Engineer: HPC & GPU Clusters

Advanced Micro Devices, Inc. • San Jose (CA)

On-site
USD 180,000 - 260,000
AI Systems Engineer: HPC & GPU Clusters
AI Systems Engineer: HPC & GPU Clusters

AMD • San Jose (CA)

On-site
USD 180,000 - 260,000
AMD benefits
AI/HPC Cluster Architect
AI/HPC Cluster Architect

AMD • Austin (TX)

On-site
USD 140,000 - 200,000
AI/HPC Cluster Architect
AI/HPC Cluster Architect

AMD • United States

On-site
USD 140,000 - 210,000
AI Systems Engineer - HPC
AI Systems Engineer - HPC

Advanced Micro Devices • San Jose (CA), Northern (KY)

Hybrid
USD 140,000 - 190,000
AI Systems Engineer - HPC
AI Systems Engineer - HPC

Advanced Micro Devices, Inc. • San Jose (CA)

On-site
USD 180,000 - 260,000
HPC Systems Engineer - AI Workloads
HPC Systems Engineer - AI Workloads

Advanced Micro Devices • San Jose (CA)

On-site
USD 190,000 - 230,000
HPC Systems Engineer - AI Workloads
HPC Systems Engineer - AI Workloads

AMD • San Jose (CA)

On-site
USD 140,000 - 210,000
AI Systems Engineer - HPC
AI Systems Engineer - HPC

AMD • San Jose (CA)

On-site
USD 180,000 - 260,000
AMD benefits
Senior Datacenter Platform Engineer — GPU/AI Infra
Senior Datacenter Platform Engineer — GPU/AI Infra

Advanced Micro Devices, Inc. • Austin (TX)

On-site
USD 110,000 - 160,000
AMD benefits at a glance