AI/HPC Cluster Architect - Scalable Data Center Design

Advanced Micro Devices, Inc.

Austin (TX)

On-site

USD 140,000 - 190,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

AMD benefits
Equal opportunity employer
Visa sponsorship not available

Job summary

Advanced Micro Devices, Inc. in Austin is seeking a highly skilled systems engineer to architect and design scalable AI/HPC clusters, evaluating compute, storage, networking, and power delivery to optimize performance across deployments.

You will collaborate with cross-functional teams to deliver cutting-edge infrastructure for AI and HPC workloads. The role requires expertise in HPC, AI infrastructure, data center systems engineering, and deep knowledge of power delivery, GPUs/CPUs,

Qualifications

  • Experience in HPC, AI infrastructure, or data center systems engineering.
  • Strong understanding of rack power delivery.
  • Knowledge of GPU/CPU architectures, PCIe, InfiniBand, and Ethernet networking.
  • Excellent problem-solving, communication, and documentation skills.

Responsibilities

  • System Architecture & Design: Design scalable AI/HPC clusters including compute, storage, and networking with specific focus on power delivery.
  • Evaluate and select CPUs, GPUs, accelerators, interconnects, and memory configurations for optimal cluster performance.
  • Power: Design leading-edge power delivery solutions for high-density AI/GPU deployments.
  • Define power budgets, redundancy schemes, and fault tolerance mechanisms.
  • Network: Design network topologies to maximize overall cluster performance.
  • Understand the network performance needs of different types of workloads.
  • Understand advantages and performance trade-offs of network topologies for AI/HPC clusters.
  • Storage: Design and optimize storage solutions to maximize AI/HPC cluster performance.
  • Understand advantages and performance trade-offs of cluster storage solutions, e.g. Lustre, Ceph, etc.
  • Collaboration: Work across multiple organizations with subject matter experts from hardware, software, network, data center, and operations teams to deliver scalable, efficient, and reliable compute infrastructure.

Skills

HPC
AI infra
Data center systems
Power delivery
GPU/CPU architectures
Networking (PCIe/IB/Ethernet)
Storage (Lustre/Ceph)
Collaboration

Education

Bachelor's or Master's in Electrical Engineering, Computer Engineering, or Computer Science

Job description

Advanced Micro Devices, Inc. in Austin is seeking a highly skilled systems engineer to architect and design scalable AI/HPC clusters, evaluating compute, storage, networking, and power delivery to optimize performance across deployments.

You will collaborate with cross-functional teams to deliver cutting-edge infrastructure for AI and HPC workloads. The role requires expertise in HPC, AI infrastructure, data center systems engineering, and deep knowledge of power delivery, GPUs/CPUs,

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI/HPC Cluster Architect: Power, Network & Systems
AI/HPC Cluster Architect: Power, Network & Systems

Advanced Micro Devices • Austin (TX)

On-site
USD 120,000 - 180,000
AMD benefits
AI/HPC Cluster Architect
AI/HPC Cluster Architect

Advanced Micro Devices • Austin (TX)

On-site
USD 120,000 - 180,000
AMD benefits
AI/HPC Cluster Architect — Data Center Power & Network
AI/HPC Cluster Architect — Data Center Power & Network

AMD • Austin (TX)

On-site
USD 120,000 - 180,000
AMD benefits at a glance
AI/HPC Cluster Architect
AI/HPC Cluster Architect

AMD • Austin (TX)

On-site
USD 140,000 - 200,000
AI/HPC Cluster Architect
AI/HPC Cluster Architect

Socket.dev • Austin (TX)

On-site
USD 140,000 - 230,000
AI/HPC Storage Architect — System Design Engineer
AI/HPC Storage Architect — System Design Engineer

AMD • Austin (TX)

On-site
USD 140,000 - 210,000
AI Systems Engineer: HPC & GPU Clusters
AI Systems Engineer: HPC & GPU Clusters

Advanced Micro Devices, Inc. • San Jose (CA)

On-site
USD 180,000 - 260,000
AI HPC Storage Architect & Reference Designs
AI HPC Storage Architect & Reference Designs

Advanced Micro Devices • Austin (TX)

On-site
USD 180,000 - 240,000
AI Cluster Storage Architect & System Designer
AI Cluster Storage Architect & System Designer

Advanced Micro Devices, Inc. • Austin (TX)

On-site
USD 150,000 - 230,000
AI Cluster & Data Center Design Engr
AI Cluster & Data Center Design Engr

AMD • Austin (TX)

On-site
USD 120,000 - 180,000
AMD benefits at a glance