AI Systems Engineer: HPC & GPU Clusters

AMD

San Jose (CA)

On-site

USD 180,000 - 260,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

AMD benefits

Job summary

AMD in San Jose, CA is seeking an AI Systems Engineer to join our IT compute platforms team. You will design, deploy, and manage HPC infrastructure, GPU clusters, and AI workload schedulers to enable scalable AI services on AMD hardware.

You should have passion for large-scale distributed computing, experience across a globally distributed organization, and a drive to deliver end-to-end outcomes with high reliability and performance.

Qualifications

  • Bachelor's or master's degree in computer science or computer engineering preferred.
  • Experience with HPC/AI workload management and distributed systems.
  • Experience across a globally distributed organization is desirable.

Responsibilities

  • Develop, implement, and maintain GPU-based clusters with optimal performance.
  • Administer ML/AI platforms – distributed ML services, LLMs and AI inference, including deployments and monitoring.
  • Automate system provisioning and end-to-end cluster management.
  • Collaborate with cross-functional teams to address AI infrastructure requirements and support AI projects.
  • Monitor performance and ensure adherence to industry best practices and company standards.
  • Use AI/ML to improve internal processes and tools used in service delivery.

Skills

Python
Distributed computing
Problem solving
Communication skills
Project management

Education

Bachelor's or Master's in CS/CE

Tools

SLURM
Kubernetes
Terraform
Prometheus
Grafana
Ansible
Saltstack

Job description

AMD in San Jose, CA is seeking an AI Systems Engineer to join our IT compute platforms team. You will design, deploy, and manage HPC infrastructure, GPU clusters, and AI workload schedulers to enable scalable AI services on AMD hardware.

You should have passion for large-scale distributed computing, experience across a globally distributed organization, and a drive to deliver end-to-end outcomes with high reliability and performance.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI Systems Engineer: HPC & GPU Clusters
AI Systems Engineer: HPC & GPU Clusters

Advanced Micro Devices, Inc. • San Jose (CA)

On-site
USD 180,000 - 260,000
AI HPC Systems Engineer: GPU Clusters & ML Platforms
AI HPC Systems Engineer: GPU Clusters & ML Platforms

AMD • San Jose (CA)

On-site
USD 140,000 - 210,000
AI Systems Engineer - HPC
AI Systems Engineer - HPC

Advanced Micro Devices, Inc. • San Jose (CA)

On-site
USD 180,000 - 260,000
AI/HPC Infrastructure Architect
AI/HPC Infrastructure Architect

AMD • Austin (TX)

On-site
USD 140,000 - 190,000
AMD benefits at a glance
AI Systems Engineer - HPC
AI Systems Engineer - HPC

AMD • San Jose (CA)

On-site
USD 180,000 - 260,000
AMD benefits
HPC Systems Engineer - AI Workloads
HPC Systems Engineer - AI Workloads

Advanced Micro Devices • San Jose (CA)

On-site
USD 190,000 - 230,000
HPC Systems Engineer - AI Workloads
HPC Systems Engineer - AI Workloads

AMD • San Jose (CA)

On-site
USD 140,000 - 210,000
Director of AI Infrastructure & HPC Connectivity
Director of AI Infrastructure & HPC Connectivity

AMD • San Jose (CA)

On-site
USD 180,000 - 230,000
AI/HPC Cluster Design Engineer
AI/HPC Cluster Design Engineer

AMD • Austin (TX)

On-site
USD 140,000 - 190,000
AMD benefits at a glance
AI/HPC Cluster Design Engineer
AI/HPC Cluster Design Engineer

Advanced Micro Devices, Inc. • Austin (TX)

On-site
USD 120,000 - 190,000
AMD benefits at a glance