AI Systems Engineer: HPC & GPU Clusters

Advanced Micro Devices, Inc.

San Jose (CA)

On-site

USD 180,000 - 260,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

AMD is seeking an AI Systems Engineer to design, deploy, and manage HPC/AI infrastructure, GPU clusters, and AI workload schedulers. You will collaborate across teams to deliver scalable, high-performance AI services on AMD hardware, with a focus on end-to-end reliability and security.

You will optimize distributed AI workloads, implement ML platforms, and advance automation for provisioning and operation of HPC resources in a globally distributed environment.

Qualifications

  • Bachelor's or Master's degree in computer science or computer engineering preferred.
  • Experience in Python-based AI apps.
  • HPC infrastructure engineering for AI/HPC.
  • SLURM and Kubernetes management.
  • GPU clusters management and optimization.
  • Web services with HPC backend (AI).
  • Proficiency in RoCEv2, K8s, KVM, Ubuntu, Python, Shell, GPU drivers, and 400G networking.
  • AI workload schedulers and allocation optimization.
  • Automation/monitoring tools: Ansible, Terraform, Prometheus, Grafana.

Responsibilities

  • Develop, implement, and maintain GPU-based clusters, ensuring optimal performance.
  • Administer ML/AI platforms - Distributed ML services, LLMs and AI inferencing, by managing deployments, resource allocation, monitoring, and security.
  • Automate system provisioning and Cluster management end to end.
  • Collaborate with cross-functional teams to address AI infrastructure requirements, support AI-related projects, and provide technical expertise.
  • Monitor and evaluate the performance of AI systems and clusters, ensuring that they adhere to industry best practices and meet company standards.
  • Use AI/ML to continuously improve internal processes and tools used in end-to-end delivery of services in this team.

Skills

Python
Distributed computing
Communication
Problem solving

Education

Bachelor's or Master's in CS/CE

Tools

SLURM
Kubernetes
Ansible
Terraform
Prometheus
Grafana
RoCEv2
KVM
Ubuntu
GPU drivers
400G networking

Job description

AMD is seeking an AI Systems Engineer to design, deploy, and manage HPC/AI infrastructure, GPU clusters, and AI workload schedulers. You will collaborate across teams to deliver scalable, high-performance AI services on AMD hardware, with a focus on end-to-end reliability and security.

You will optimize distributed AI workloads, implement ML platforms, and advance automation for provisioning and operation of HPC resources in a globally distributed environment.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI HPC Systems Engineer: GPU Clusters & ML Platforms
AI HPC Systems Engineer: GPU Clusters & ML Platforms

AMD • San Jose (CA)

On-site
USD 140,000 - 210,000
AI Systems Engineer: HPC & GPU Clusters
AI Systems Engineer: HPC & GPU Clusters

AMD • San Jose (CA)

On-site
USD 180,000 - 260,000
AMD benefits
AI/HPC Infrastructure Architect
AI/HPC Infrastructure Architect

AMD • Austin (TX)

On-site
USD 140,000 - 190,000
AMD benefits at a glance
AI Systems Engineer - HPC
AI Systems Engineer - HPC

Advanced Micro Devices, Inc. • San Jose (CA)

On-site
USD 180,000 - 260,000
HPC Systems Engineer - AI Workloads
HPC Systems Engineer - AI Workloads

Advanced Micro Devices • San Jose (CA)

On-site
USD 190,000 - 230,000
AI Systems Engineer - HPC
AI Systems Engineer - HPC

AMD • San Jose (CA)

On-site
USD 180,000 - 260,000
AMD benefits
HPC Systems Engineer - AI Workloads
HPC Systems Engineer - AI Workloads

AMD • San Jose (CA)

On-site
USD 140,000 - 210,000
AI/HPC Cluster Design Engineer
AI/HPC Cluster Design Engineer

AMD • Austin (TX)

On-site
USD 140,000 - 190,000
AMD benefits at a glance
AI/HPC Cluster Design Engineer
AI/HPC Cluster Design Engineer

Advanced Micro Devices, Inc. • Austin (TX)

On-site
USD 120,000 - 190,000
AMD benefits at a glance
Senior Datacenter Platform Engineer — GPU/AI Infra
Senior Datacenter Platform Engineer — GPU/AI Infra

Advanced Micro Devices, Inc. • Austin (TX)

On-site
USD 110,000 - 160,000
AMD benefits at a glance