AI Systems Engineer: HPC & GPU Clusters

Advanced Micro Devices, Inc.

San Jose (CA)

On-site

USD 180,000 - 260,000

Full time

4 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

AMD is seeking an AI Systems Engineer to design, deploy, and manage HPC/AI infrastructure, GPU clusters, and AI workload schedulers. You will collaborate across teams to deliver scalable, high-performance AI services on AMD hardware, with a focus on end-to-end reliability and security.

You will optimize distributed AI workloads, implement ML platforms, and advance automation for provisioning and operation of HPC resources in a globally distributed environment.

Qualifications

  • Bachelor's or Master's degree in computer science or computer engineering preferred.
  • Experience in Python-based AI apps.
  • HPC infrastructure engineering for AI/HPC.
  • SLURM and Kubernetes management.
  • GPU clusters management and optimization.
  • Web services with HPC backend (AI).
  • Proficiency in RoCEv2, K8s, KVM, Ubuntu, Python, Shell, GPU drivers, and 400G networking.
  • AI workload schedulers and allocation optimization.
  • Automation/monitoring tools: Ansible, Terraform, Prometheus, Grafana.

Responsibilities

  • Develop, implement, and maintain GPU-based clusters, ensuring optimal performance.
  • Administer ML/AI platforms - Distributed ML services, LLMs and AI inferencing, by managing deployments, resource allocation, monitoring, and security.
  • Automate system provisioning and Cluster management end to end.
  • Collaborate with cross-functional teams to address AI infrastructure requirements, support AI-related projects, and provide technical expertise.
  • Monitor and evaluate the performance of AI systems and clusters, ensuring that they adhere to industry best practices and meet company standards.
  • Use AI/ML to continuously improve internal processes and tools used in end-to-end delivery of services in this team.

Skills

Python
Distributed computing
Communication
Problem solving

Education

Bachelor's or Master's in CS/CE

Tools

SLURM
Kubernetes
Ansible
Terraform
Prometheus
Grafana
RoCEv2
KVM
Ubuntu
GPU drivers
400G networking

Job description

AMD is seeking an AI Systems Engineer to design, deploy, and manage HPC/AI infrastructure, GPU clusters, and AI workload schedulers. You will collaborate across teams to deliver scalable, high-performance AI services on AMD hardware, with a focus on end-to-end reliability and security.

You will optimize distributed AI workloads, implement ML platforms, and advance automation for provisioning and operation of HPC resources in a globally distributed environment.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Systems Engineer: HPC & GPU Clusters
AI Systems Engineer: HPC & GPU Clusters

AMD • San Jose (CA)

On-site
USD 180,000 - 260,000
AMD benefits
AI/HPC Cluster Architect
AI/HPC Cluster Architect

AMD • Austin (TX)

On-site
USD 140,000 - 200,000
AI/HPC Cluster Architect
AI/HPC Cluster Architect

Socket.dev • Austin (TX)

On-site
USD 140,000 - 230,000
AI Systems Engineer - HPC
AI Systems Engineer - HPC

Advanced Micro Devices, Inc. • San Jose (CA)

On-site
USD 180,000 - 260,000
AI Systems Engineer - HPC
AI Systems Engineer - HPC

Advanced Micro Devices • San Jose (CA), Northern (KY)

Hybrid
USD 140,000 - 190,000
AI Systems Engineer - HPC
AI Systems Engineer - HPC

AMD • San Jose (CA)

On-site
USD 180,000 - 260,000
AMD benefits
AI Systems Engineer: ML Kernels & HPC Acceleration
AI Systems Engineer: ML Kernels & HPC Acceleration

Socket.dev • San Jose (CA)

Hybrid
USD 150,000 - 190,000
AI/HPC Cluster Architect — Data Center Power & Network
AI/HPC Cluster Architect — Data Center Power & Network

AMD • Austin (TX)

On-site
USD 120,000 - 180,000
AMD benefits at a glance
AI Cluster & Data Center Design Engr
AI Cluster & Data Center Design Engr

Advanced Micro Devices, Inc. • Austin (TX)

On-site
USD 140,000 - 190,000
AMD benefits
Equal opportunity employer
Visa sponsorship not available
AI/HPC Cluster Design Engineer
AI/HPC Cluster Design Engineer

Advanced Micro Devices, Inc. • Austin (TX)

On-site
USD 120,000 - 190,000
AMD benefits at a glance