AI Systems Engineer: HPC & GPU Clusters

AMD

San Jose (CA)

On-site

USD 180,000 - 260,000

Full time

3 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

AMD benefits

Job summary

AMD in San Jose, CA is seeking an AI Systems Engineer to join our IT compute platforms team. You will design, deploy, and manage HPC infrastructure, GPU clusters, and AI workload schedulers to enable scalable AI services on AMD hardware.

You should have passion for large-scale distributed computing, experience across a globally distributed organization, and a drive to deliver end-to-end outcomes with high reliability and performance.

Qualifications

  • Bachelor's or master's degree in computer science or computer engineering preferred.
  • Experience with HPC/AI workload management and distributed systems.
  • Experience across a globally distributed organization is desirable.

Responsibilities

  • Develop, implement, and maintain GPU-based clusters with optimal performance.
  • Administer ML/AI platforms – distributed ML services, LLMs and AI inference, including deployments and monitoring.
  • Automate system provisioning and end-to-end cluster management.
  • Collaborate with cross-functional teams to address AI infrastructure requirements and support AI projects.
  • Monitor performance and ensure adherence to industry best practices and company standards.
  • Use AI/ML to improve internal processes and tools used in service delivery.

Skills

Python
Distributed computing
Problem solving
Communication skills
Project management

Education

Bachelor's or Master's in CS/CE

Tools

SLURM
Kubernetes
Terraform
Prometheus
Grafana
Ansible
Saltstack

Job description

AMD in San Jose, CA is seeking an AI Systems Engineer to join our IT compute platforms team. You will design, deploy, and manage HPC infrastructure, GPU clusters, and AI workload schedulers to enable scalable AI services on AMD hardware.

You should have passion for large-scale distributed computing, experience across a globally distributed organization, and a drive to deliver end-to-end outcomes with high reliability and performance.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI/HPC Cluster Architect
AI/HPC Cluster Architect

Socket.dev • Austin (TX)

On-site
USD 140,000 - 230,000
AI/HPC Cluster Architect
AI/HPC Cluster Architect

AMD • Austin (TX)

On-site
USD 140,000 - 200,000
AI Systems Engineer - HPC
AI Systems Engineer - HPC

Advanced Micro Devices • San Jose (CA), Northern (KY)

Hybrid
USD 140,000 - 190,000
Senior System Design & AI Cluster Engineer
Senior System Design & AI Cluster Engineer

AMD • Santa Clara (CA)

On-site
USD 170,000 - 250,000
AMD Benefits at a glance
AI Systems Engineer - HPC
AI Systems Engineer - HPC

AMD • San Jose (CA)

On-site
USD 180,000 - 260,000
AMD benefits
AI Systems Engineer: ML Kernels & HPC Acceleration
AI Systems Engineer: ML Kernels & HPC Acceleration

Socket.dev • San Jose (CA)

Hybrid
USD 150,000 - 190,000
Lead AI Systems Engineer: GPU Clusters & AI Ops
Lead AI Systems Engineer: GPU Clusters & AI Ops

Semiconductor Engineering • San Jose (CA)

On-site
USD 140,000 - 220,000
Director of AI Infrastructure & HPC Connectivity
Director of AI Infrastructure & HPC Connectivity

AMD • San Jose (CA)

On-site
USD 180,000 - 230,000
AI/HPC Cluster Architect — Data Center Power & Network
AI/HPC Cluster Architect — Data Center Power & Network

AMD • Austin (TX)

On-site
USD 120,000 - 180,000
AMD benefits at a glance
AI/HPC Cluster Design Engineer
AI/HPC Cluster Design Engineer

Socket.dev • Austin (TX)

On-site
USD 140,000 - 230,000